Small Model Trained on Its Own Mistakes Reaches 80% on HumanEval, Beats GPT-3.5 on Math
WHY IT MATTERS
A reported experiment on r/LocalLLaMA describes training a small language model iteratively on its own errors, achieving 80% on HumanEval and outperforming GPT-3.5 on math benchmarks. No paper or external verification link is provided. The methodology aligns with self-improvement and rejection sampling fine-tuning techniques.
What Happened
A user on r/LocalLLaMA reported training a small language model iteratively on its own errors, claiming the resulting checkpoint scores 80% on HumanEval and outperforms GPT-3.5 on unspecified math benchmarks. The post includes no paper, model weights, base architecture, dataset composition, parameter count, or training compute disclosure. No external verification or replication has been produced.
Why It Matters
The claimed 80% HumanEval figure sits above commonly cited pass@1 scores for GPT-3.5 (roughly 48%) and GPT-4 (near 67% on the standard evaluation), which — if reproducible — would place a locally-run small model above commercial baselines on a coding benchmark that has historically tracked general reasoning capability. The methodology described maps to rejection sampling fine-tuning and self-improvement loops, both of which are established in published literature and computationally inexpensive relative to pretraining. For operators weighing local inference against API dependence, the result implies a path to competitive coding performance without per-token cost or data egress. The absence of verification means the practical takeaway is directional, not actionable: it tells builders where to look, not what to deploy. The strategic weight of the claim rests less on the specific checkpoint than on whether verifier-gated self-training can consistently close the gap between small open models and hosted frontier APIs in narrow, gradeable domains.
Technical Details
The described loop generates outputs from a base model, filters incorrect samples against a ground-truth signal (unit tests for code, answer keys for math), and retrains on corrected or passing trajectories. Rejection sampling fine-tuning of this shape typically requires a verifiable reward, which limits applicability to domains with automated graders — code execution, math problems with known answers, and structured tasks. Reported performance is on HumanEval pass@1, a 164-problem benchmark that is small enough for variance to matter and is often contaminated in public training corpora. No information is available on model size, tokenizer, context length, quantization, or whether the training set overlapped with HumanEval problems. The math claim against GPT-3.5 is stated without benchmark names, sample counts, or evaluation protocol. Standard self-training loops also risk mode collapse and reward hacking when the verifier is imperfect — a concern the post does not address.
Operational Impact
If the approach replicates, the cost structure for domain-specific coding assistants shifts: a fine-tuning pass over self-generated correct samples is orders of magnitude cheaper than pretraining and can run on modest single-node hardware. Builders currently routing coding queries to hosted APIs for quality reasons would gain a local fallback for a subset of tasks — particularly function-level synthesis with clear test harnesses. The workflow change is procedural rather than architectural: stand up an execution sandbox to score generations, curate passing samples, and schedule periodic fine-tune cycles. Teams without automated verification for their target task cannot use the loop as described. Existing evaluation pipelines remain the bottleneck — the method's ceiling is set by the quality of the reward signal, not the training compute. Operators should also budget for regression testing on held-out sets, since self-trained checkpoints can degrade on distributions outside the verifier's coverage.
SOURCE
SHARE
MORE FROM STUFFINSIDER