Radical Geek field guide

Radical Geek's Guide to Evaluating Coding Agents And Model Routing

A practical evaluation system for coding-agent outcomes, routing decisions, escalation value, tail economics and delivery impact.

For
Engineering leaders, platform teams and developers comparing agent workflows and model lanes
Format
Evaluation framework and practical harness
Access
Free PDF guide
Cover of Radical Geek's Guide to Evaluating Coding Agents And Model Routing

Why this guide

Coding-agent and model decisions need evidence from the operating conditions in which they will be used. A convincing demonstration says little about patch quality, review load, routing accuracy, escalation value, recovery or the cost of difficult work.

This guide establishes a comparative evaluation and qualification layer for real work classes. It keeps one work-item identity across retries, repairs, route changes and human review, then connects the result to accepted engineering outcomes.

Use it to decide which model, workflow and escalation policy belongs on each bounded class of work, and when the evidence supports promotion, narrowing or retirement.

Measure accepted engineering outcomes. Treat routing as a policy you can test, explain and change.

What you will leave with

A guide built to be used.

  1. 01

    Build representative evaluations from real repository work, exceptions and accepted outcomes.

  2. 02

    Measure patch quality, verification evidence, human review, latency and full delivery cost.

  3. 03

    Test whether routing and escalation policies send each work class to an appropriate model lane.

  4. 04

    Compare local, cloud, MoE and fusion workflows without transferring wider decision authority.

  5. 05

    Use qualification evidence to promote, narrow, demote or retire a model or workflow route.

Inside the guide

The working ground it covers.

  • What to measure and the evaluation levels
  • Benchmark types and a coding-agent harness
  • Patch quality and verification metrics
  • Routing, escalation and tail economics
  • Local, cloud, MoE and fusion evaluation
  • Human review and ground-truth confidence
  • A minimum dashboard and first 30 days plan