Canonical SWE-bench Lite Benchmark
Full 300-task canonical execution benchmark evaluating Qwen2.5-Coder-14B-Instruct (4-bit) running on Apple Silicon via AtacamaODR's Structural Context Compaction engine. Rather than transmitting hundreds of thousands of raw repository tokens across public cloud APIs, AtacamaODR compacts multi-file repository hierarchies on-device into high-density interface outlines, isolating the culprit fault target in seconds. Zero synthetic mock data. Zero simulated pass rates. 100% physically executed against canonical ground-truth commits.
| Target Repository | Evaluated Tasks | Accurately Localized | Localization Accuracy | Status |
|---|---|---|---|---|
| astropy/astropy | 6 | 6 | 100.0% | Verified |
| django/django | 114 | 105 | 92.1% | Verified |
| matplotlib/matplotlib | 23 | 22 | 95.7% | Verified |
| mwaskom/seaborn | 4 | 3 | 75.0% | Verified |
| pallets/flask | 3 | 3 | 100.0% | Verified |
| psf/requests | 6 | 3 | 50.0% | Verified |
| pydata/xarray | 5 | 4 | 80.0% | Verified |
| pylint-dev/pylint | 6 | 6 | 100.0% | Verified |
| pytest-dev/pytest | 17 | 16 | 94.1% | Verified |
| scikit-learn/scikit-learn | 23 | 22 | 95.7% | Verified |
| sphinx-doc/sphinx | 16 | 14 | 87.5% | Verified |
| sympy/sympy | 77 | 65 | 84.4% | Verified |
| Consolidated Total (12 Repositories) | 300 | 269 | 89.7% | Physically Verified |
Structural Context Compaction
In repository-scale benchmarks like SWE-bench Lite, large language models are overwhelmed by raw context windows exceeding 500,000 lines across hundreds of files. AtacamaODR compacts repository structures directly on-device into compact interface outlines. This preserves complete functional contract definitions while discarding internal implementation noise, allowing the 14B on-device model to evaluate repository architecture in sub-4-second horizons.
Economics of On-Device Localization
Frontier AI agents typically spend $0.50 to $2.00 in cloud prompt tokens simply reading repository directory structures to locate where an issue exists. By resolving 89.7% of fault localization targets on-device at $0.00 cost, AtacamaODR acts as an autonomous pre-flight filter, reserving expensive frontier cloud escalation solely for complex bounded code patch generation.
This evaluation was executed against the official canonical SWE-bench Lite split (300 tasks). For each task, the problem statement was provided along with the compacted repository interface spine. The model predicted the target modification path, which was programmatically evaluated against the ground-truth commit diff from the official SWE-bench repository. Zero synthetic mock scripts or synthetic rate simulators were utilized.