bare qwen3:14b passes 79.9% of HumanEval+ (all 164 problems), running entirely on a dedicated NVIDIA DGX Spark — no cloud, no metered tokens, your code never leaving the box. Chad runs the model's code, checks it against tests, and has the model fix what fails. We preload that setup and rent you the whole machine for $100 a week. The first 24 hours are free.
A note on our with-Chad benchmark numbers: we found a scoring issue and pulled them while we re-measure honestly. The bare baselines stand. Here's exactly what happened.
Every row is a full 164-problem HumanEval+ run on the same class of DGX Spark you'd be renting — one attempt per problem, no feedback. These are the verified bare baselines. The with-Chad comparison is being re-measured under a stricter rule after we caught a scoring issue; the full story and methodology are on the results page.
| Model | bare HumanEval+ |
|---|---|
| qwen2.5-coder:14b | 82.9% |
| qwen3:14b | 79.9% |
| gemma3:27b | 75.6% |
| nemotron:70b | 73.2% |
| codestral:22b | 73.2% |
| mistral-nemo:12b | 52.4% |
Bare baselines, verified. Honest with-Chad numbers get added to the Spark Quality Index as the re-runs complete.
One flat price. No tokens, no usage caps, no surprise bill at the end of the month.
$100 / week · week to week · stop any time
NVIDIA DGX Spark (GX10, 128 GB unified memory). Not a shared slice, not a queue, not a spot instance that disappears. Yours for the week over SSH.
qwen3:14b served locally with Chad configured and Aider wired up and ready. Open your repo and start.
The model runs on the box. Prompts, code, and data stay local. Nothing is sent to a cloud API, because there is no cloud API.
It's your machine. The tuned coding setup is the default, but you can pull other models, run your own stack, or use it as a private inference box.
Direct help getting your project running, and Chad updates as we ship them. No ticket queue.
The same hardened runner that produces our published numbers keeps the box healthy: watchdogs, auto-recovery, and receipts for what ran.
Before you spend a dollar, spend a day on the machine — real SSH access, the full setup, your own code. Eval slots are limited because each one occupies dedicated hardware, so we screen applications. Answer the six questions below; if it looks like a fit, you'll get credentials and a start window within a day.
Nothing automatic. There's no card on file and nothing renews itself. If you want the machine, reply to the email and pay for your first week. If not, we wipe your data from the box and that's that.
Yes. That's exactly why we screen — each eval ties up a dedicated machine for a day, so the slots go to people genuinely deciding whether to rent.
Nowhere. The model runs on the machine you're renting. When your rental ends, your files are deleted. We don't train on your code, and we couldn't even if we wanted to — it never reaches us.
Yes. The machine ships with the tuned qwen3:14b + Chad setup because that's the configuration we publish numbers for, but the Ollama model library is a command away and the 128 GB of unified memory fits models far larger than 14B.
An NVIDIA DGX Spark (GX10): Blackwell GPU, 128 GB unified memory, aarch64. The same class of machine every number on our results page was measured on.
Ask. Weekly is the default; longer commitments get a better rate.