This page documents how the OpenCode, Cline and Command Code comparison published in OpenCode vs Cline vs Command Code: Which Costs Less? was run. The goal is for anyone to be able to repeat the test, with another model or other harnesses, and get comparable numbers.
The point is to isolate the harness. We pin the model, the tasks, the initial repository state, the prompt and the permission mode; the only thing that changes is who orchestrates. The model goes into the log as a column, not as an axis, so the same battery can run with a different model three months from now and the numbers stay comparable.
Environment: Windows 11, no Docker. Harnesses: OpenCode 2.0.14, Cline 3.0.62 and Command Code 1.66.0. Model: Space Bunny (stealth preview, September 2026).
Some traps along this path produce numbers that look publishable and are wrong. Four of them showed up in our first run and are instrument errors, meaning they come from how we call the harness. A fifth showed up in the Pixel Canary test and is different: it’s a property of the harness itself. They are all documented here because they are the most reusable part of this methodology.
Results from every test run with this methodology are collected on the Stealth Models Benchmark page.
Updated October 1, 2026, with the Fledge Alpha test: look at the per-run cache distribution and not just the median (section 4), connection drops recorded separately, a rule for rerun rounds and a fix to the seed (leftover .git/objects). On September 27, 2026: trap 3.4 (cache only on the system prompt), an axis variant for models available in a single harness, and a new checklist item.
1. Experiment design
The isolated variable
Pinned across every run: model, task set, literal prompt, initial repository state, permission mode (full auto-approval), reasoning effort (provider default) and compaction mode.
What the harness does differently (system prompt, tool definitions, file-reading policy, context compaction) isn’t noise. It’s the treatment, and it’s exactly what we want to see show up.
Confounders and what to do about each
Route to the model. Each harness reaches the model through its own gateway, so latency and quota differ by construction. Without a unified bring-your-own-key setup you can’t eliminate this; you can declare it and avoid drawing conclusions about time. If you have a shared provider available (OpenRouter or similar) and it keeps up with the pace, unifying the route is the cleanest design.
Harness version. CLIs update themselves and change the model catalog under your feet. Pin the version, record it on every result row and disable auto-update during the run.
Run order. Don’t run 15 of one and then 15 of another: API load over the day turns into a harness difference. Shuffle with a fixed seed.
When the model exists in only one harness
A stealth model sometimes shows up in only one harness’s catalog. Then you can’t pin the model and vary the harness. The way out is to flip the axis: pin the harness and vary the model, comparing the new model with another one that has already run on the same harness. That’s what we did with Pixel Canary, which only existed in Command Code, and with Fledge Alpha, which only existed in OpenCode.
That result doesn’t mix with the main comparison. The harness’s way of spending context and using cache is baked into both sides, so every table has to say which axis varied. Record availability in each catalog (harnesses and OpenRouter) with the date of the check, and run the full battery if the model later shows up in other harnesses.
Isolation without Docker
Containers in public benchmarks exist for reproducibility across machines. To compare harnesses on the same machine, strict state reset is enough:
- An immutable
seed/directory. Each run copies the seed into a throwaway directory. - Reset by copying, never by
git checkout. Checkout doesn’t remove untracked files the agent created. - Remove
.gitfrom the seed. With history, the agent finds the injected bug from the diff instead of reading the code. - Check that nothing of
.gitis left behind. In our first three batteries, the seed still had leftover.git/objects(no HEAD or refs, unrecognizable to git) from the setup. It applied equally to every model and harness, so it doesn’t change the comparisons, but it has been fixed for future runs.
Clean profiles
Each harness runs with HOME and XDG_CONFIG_HOME pointing to a throwaway profile that contains only the credential. No MCP, no skills, no plugins, no agents, no AGENTS.md or CLAUDE.md, and no learned-preference files (Command Code keeps a taste/ directory that has to stay out). The user’s real configuration is never touched.
Check in the dry run that none of the harnesses loaded anything, and record that check.
Tasks
Bugs were injected into different modules of a real public repository, chosen so the symptom sits far from the cause. That’s what forces the agent to read the code instead of solving it with a grep.
Avoid “implement feature X that exists upstream” tasks: in a well-known repository the model answers from memory and the harness never has to manage context. An injected bug defeats memorization and gives you an evaluation suite for free.
We chose pallets/jinja 3.2.0.dev: pure Python, a single dependency, 911 tests that run in 4 seconds on Windows, 14k lines across 25 files, BSD-3 license. The fast suite matters: with 45 runs, slow grading would dominate the clock. The five bugs went into sandbox.py, lexer.py, runtime.py, loaders.py and compiler.py.
Visible tests, grading against a pristine copy
We left tests/ visible to the agent. Hiding it (SWE-bench style) would remove the loop of running a test and watching it fail, and that loop is one of the biggest differences between harnesses.
Integrity comes from grading: before evaluation, tests/ is replaced with a pristine copy and the suite runs against it. Tampering with tests doesn’t pass, and the earlier diff of tests/ becomes a cheating detector.
Execution
Three repetitions per task × harness pair, order shuffled with a fixed seed: 45 runs, 22,957,016 tokens processed in 1.5 hours of execution.
2. Metrics
Outcome. Pass or fail on the suite, per task. A timeout is a separate outcome from “not solved”: we had a run that timed out with the fix already written. A connection drop at the end of the session (ECONNRESET) follows the same logic: if the fix is already on disk and the suite passes, the run counts as solved, and the drop is recorded separately as gateway instability.
Cost and efficiency. Uncached input, cache reads, cache writes, output, reasoning, turns, tool calls and time.
Harness-specific, which is what justifies the whole exercise:
- Scaffolding: context in the first request, before any work. System prompt plus tool definitions. It’s the fixed tax per task.
- Total context: uncached input plus cache reads.
- Cache rate: cache reads divided by total context.
- Turn-by-turn cache:
inputTokensversuscacheReadTokenson each turn of a run. Shows whether the cache follows the conversation or is stuck on a fixed block (see trap 3.4). - Tool calls and turns, to read the strategy (slice a lot or pull everything at once).
- Malformed tool calls and edit retries: with the model fixed, this measures the quality of the interface the harness gives the model.
Behavior. Writes outside the working directory, modified tests, deleted files. Any failure here invalidates the run.
Converting to money. Tokens were priced at DeepSeek V4.1 Flash rates, with a separate price per category: $0.15 per million input, $0.60 per million output and $0.003 per million cache reads.
3. The traps
3.1 The .cmd shim truncates the prompt (Windows)
npm-installed CLIs on Windows expose a .cmd shim. Calling that shim goes through cmd.exe, which cuts the command line at the first line break in the argument and throws away the rest.
A multi-line prompt passed to a .cmd shim means every flag after it disappears. In our case that dropped -m: the run used the default model instead of the model under test, finished the task successfully, returned exit code 0 and would have gone into the table as a legitimate result.
Call the real executables, never the shims:
opencode.exe
cline/node_modules/@cline/cli-windows-x64/bin/cline.exe
node command-code/dist/index.mjs
The safeguard that came out of this: on every run, record which model the harness says it used, pulled from the event stream, instead of trusting the flag you passed. Not every harness exposes this; when it doesn’t, declare the limitation.
3.2 text=True decodes as cp1252
In Python on Windows, subprocess.run(..., text=True) decodes using the locale encoding. Agent output is UTF-8. The reader thread blows up with UnicodeDecodeError, stdout becomes None, and the run is lost in a way that looks like a harness error.
Always use encoding="utf-8", errors="replace".
3.3 Raw input tokens are misleading
The metric every comparison publishes is the most misleading one here. Between OpenCode and Cline, “input tokens” showed a gap of about 9x. Once cache was added in, both sent similar context, and the real difference was the cache rate: 92% versus 47%.
Raw input vs. total context
Tokens per task, median. Raw input suggests a 9x gap; total context is similar.
Uncached input (full price)
Total context (input + cache reads)
Cache reads cost a fraction of fresh input (on DeepSeek V4.1 Flash, $0.003 per million versus $0.15, or 50 times less). Report total context and cache rate as primary metrics, and convert to money with separate prices per category. I explain the concept in more detail in the post on what a cache hit is.
3.4 cache_control on the system prompt only skews the cache rate
A harness can report a low cache rate (17% to 48% in our case) even when the provider supports true incremental caching. The cause may be the harness itself, not the model or the provider.
In Command Code, the function that builds the system message (toWireSystem) only marks cache_control: {type: "ephemeral"} on two fixed blocks: the instructions and the tool definitions. Turn history, tool call results and everything that grows during the task never get that marker. In practice, the cache hit ceiling in this harness is the size of that fixed block relative to total context. It isn’t incremental caching.
Turn-by-turn cache: Pixel Canary on the lexer task
Tokens per turn, on Command Code. Input grows with the conversation; cache stays stuck on a few fixed values.
Turn 1
Turn 2
Turn 3
Turn 4
Turn 5
Turn 6
With two models on the same Command Code (Space Bunny at 48%, Pixel Canary at 17%) and the same Space Bunny reaching 92% on OpenCode, the ceiling shows: the limitation belongs to the harness, not the model. The gap from 48% to 17% inside Command Code most likely comes from how well each model’s backend honors the cache directive the harness offers. It’s residual variation on top of a ceiling that is already low by design.
How to check: look token by token inside a run (inputTokens versus cacheReadTokens per turn, in the raw log). If cache_read doesn’t grow along with the conversation history, and instead jumps between a few fixed values while input climbs, the cache is limited to a static block rather than the whole prefix. Don’t assume that a cache rate advertised by a harness vendor applies to the full conversation: it may refer only to the fixed part, or to a better scenario than the observed behavior.
It isn’t a test rig bug. Unlike the other traps in this section, which come from how we call the harness, this is a real property of the harness in normal use. Every user of the tool hits the same limitation. Don’t “fix” it by patching the installed package to match the other harnesses: that would measure a version nobody actually runs and destroy exactly what the comparison is supposed to reveal.
3.5 __pycache__ becomes a false cheating positive
pytest creates __pycache__ inside tests/ when it runs. A tampering detector that compares whole directories flags every run that executed the suite. Compare source files only.
General safeguard: keep the raw log
Our first version deleted the run directory, log included, right after grading. Without the raw log there’s no diagnosis afterward. Persist the full output of every run outside the throwaway directory.
4. Statistics and logging
At least 3 repetitions per task × harness pair, 5 if the budget allows. Report median and spread, never the mean: one run that hit the timeout skews the mean and doesn’t skew the median.
Look at the distribution, not just the median. In the Fledge Alpha test, the median cache hit was 91%, practically the same as Space Bunny’s (92%). But run by run the cache was bimodal: 11 runs between 88% and 97% and 4 runs between 9% and 28%. The bad runs cost 3.3 times more than the good ones. Whenever a metric weighs on cost, also report the per-run range and count how many runs land in each group.
Cache hit per run: the range of each group
From the lowest to the highest cache rate among the runs in the group, on a 0 to 100% scale.
Declare rerun rounds. If a run fails because of the runner, not the model or the harness, redo it and record that it was rerun and when. On Fledge Alpha, one run failed during directory cleanup and was redone two hours after the other 14.
The main metric isn’t success rate, it’s cost and tokens per solved task. With converged agents, success ties and efficiency separates them.
One JSONL line per run, fixed schema, with model, harness, harness_version, task, rep, timestamp and every metric. The model as a column is what makes the benchmark reusable: run it with another model, append to the same file, and the comparison is ready.
5. Checklist before publishing any number
- Real executables, not
.cmdshims encoding="utf-8"on every subprocess capture- Model reported by the harness recorded and checked on every run
- Raw log persisted outside the run directory
.gitremoved from the seed, with no leftovers (not even.git/objects)- Clean profiles verified in the dry run
- Grading done with pristine
tests/restored - Cheating detector ignoring execution artifacts
- Total context and cache rate, not just input tokens
- Timeout recorded separately from “not solved”
- Cache rate checked token by token (does
cache_readgrow with the conversation or stay stuck on a fixed block?) - Per-run cache distribution checked, not just the median
- Rerun rounds and connection drops declared
- Run order shuffled
- Routes to the model declared, and time read in light of them
Do a dry run of one task on every harness before the full battery. Our first dry run broke in four places, all in the instrument and none in the harnesses. Three of them would have produced a publishable, wrong table.