Spark Quality Index

What local models actually score on real local hardware. Every number below is a full 164-problem HumanEval+ run executed on an NVIDIA DGX Spark (GX10, 128 GB unified memory) — the same class of machine you can rent from us. No cloud numbers passed off as local performance, no cherry-picked subsets, and the runs that broke are listed right here instead of quietly deleted.

bare is the model alone, one attempt per problem — those numbers are verified and published below. The with-Chad comparison is being re-measured right now under a stricter rule; the next section explains exactly why and what changed.

Re-verification in progress

The with-Chad numbers are being re-measured

We hold this page to the standard in the footer, so here is the honest state of it. While reviewing how the with-Chad runs were scored, we found that Chad's repair step was allowed to read the benchmark's own grading tests when it retried a problem. That inflates a benchmark score and does not reflect what the tool does for you on your own code, so every previously published with-Chad and lift figure has been pulled until the runs are redone under a stricter rule: on retry, Chad may only see the examples in the problem statement, never the hidden grading tests.

The bare numbers below are unaffected — a bare run takes one attempt with no feedback of any kind — and they stand. New with-Chad numbers go back up here the moment the honest re-runs land. If you want to be told when they do, ask.
Bare baselines — verified

What each model scores on its own, on this hardware

One attempt per problem, no feedback, no retries. These are the honest floor: what the raw model does on a DGX Spark before any tooling. The with-Chad rows return here once the re-runs described above are complete.

RunPassRateAvg triesWall timeDate
bare qwen2.5-coder:14b136/16482.9%1.0047m2026-05-23
bare qwen3:14b131/16479.9%1.00283m2026-05-23
bare gemma3:27b124/16475.6%1.0096m2026-05-23
bare codestral:22b120/16473.2%1.0041m2026-05-23
bare nemotron:70b120/16473.2%1.00144m2026-07-08
bare mistral-nemo:12b86/16452.4%1.0020m2026-07-05

HumanEval+ bare baselines. The with-Chad comparison is being re-measured under the stricter retry rule described above and will be republished when the honest runs land.

What we don't publish

Excluded and quarantined runs

A results page you can trust has to show its trash can, not just its trophies.

Quarantined: deepseek-coder-v2:16b (both bare and Chad, 2026-05-23) — 0/164 in under six minutes, which is the signature of a broken run, not a real score. It stays out of the tables until a clean re-run replaces it.

Excluded: six partial and smoke-test runs (3–10 problems each), used for development sanity checks. Small samples flatter whoever ran them; only complete 164-problem runs make this page.
Methodology

How these numbers are made

Last updated 2026-07-08. New completed runs are added as they land. Want a model or quantization measured? Ask.