What local models actually score on real local hardware. Every number below is a full 164-problem HumanEval+ run executed on an NVIDIA DGX Spark (GX10, 128 GB unified memory) — the same class of machine you can rent from us. No cloud numbers passed off as local performance, no cherry-picked subsets, and the runs that broke are listed right here instead of quietly deleted.
bare is the model alone, one attempt per problem — those numbers are verified and published below. The with-Chad comparison is being re-measured right now under a stricter rule; the next section explains exactly why and what changed.
One attempt per problem, no feedback, no retries. These are the honest floor: what the raw model does on a DGX Spark before any tooling. The with-Chad rows return here once the re-runs described above are complete.
| Run | Pass | Rate | Avg tries | Wall time | Date |
|---|---|---|---|---|---|
| bare qwen2.5-coder:14b | 136/164 | 82.9% | 1.00 | 47m | 2026-05-23 |
| bare qwen3:14b | 131/164 | 79.9% | 1.00 | 283m | 2026-05-23 |
| bare gemma3:27b | 124/164 | 75.6% | 1.00 | 96m | 2026-05-23 |
| bare codestral:22b | 120/164 | 73.2% | 1.00 | 41m | 2026-05-23 |
| bare nemotron:70b | 120/164 | 73.2% | 1.00 | 144m | 2026-07-08 |
| bare mistral-nemo:12b | 86/164 | 52.4% | 1.00 | 20m | 2026-07-05 |
HumanEval+ bare baselines. The with-Chad comparison is being re-measured under the stricter retry rule described above and will be republished when the honest runs land.
A results page you can trust has to show its trash can, not just its trophies.
Last updated 2026-07-08. New completed runs are added as they land. Want a model or quantization measured? Ask.