AtacamaODR Apple Silicon
BENCHMARK AUDIT: TRACK 05 • CANONICAL SWE-BENCH LITE (300 TASKS) • AUTHOR: GANESH NALLASIVAM
Download Track 05 Report (.PDF)

Canonical SWE-bench Lite Benchmark

Full 300-task canonical execution benchmark evaluating Qwen2.5-Coder-14B-Instruct (4-bit) running on Apple Silicon via AtacamaODR's Structural Context Compaction engine. Rather than transmitting hundreds of thousands of raw repository tokens across public cloud APIs, AtacamaODR compacts multi-file repository hierarchies on-device into high-density interface outlines, isolating the culprit fault target in seconds. Zero synthetic mock data. Zero simulated pass rates. 100% physically executed against canonical ground-truth commits.

Fault Localization Accuracy
89.7%
269 of 300 targets pinpointed.
Total Tasks Evaluated
300 / 300
100% canonical SWE-bench Lite.
Average Turnaround Latency
4.11 s
Total suite runtime: 20.5 min.
Total Cloud API Spend
$0.00
100% on-device Metal GPU.
Repository-Level Fault Localization Ledger (300 Tasks)
Empirical Ground-Truth Verification
Target Repository Evaluated Tasks Accurately Localized Localization Accuracy Status
astropy/astropy 6 6 100.0% Verified
django/django 114 105 92.1% Verified
matplotlib/matplotlib 23 22 95.7% Verified
mwaskom/seaborn 4 3 75.0% Verified
pallets/flask 3 3 100.0% Verified
psf/requests 6 3 50.0% Verified
pydata/xarray 5 4 80.0% Verified
pylint-dev/pylint 6 6 100.0% Verified
pytest-dev/pytest 17 16 94.1% Verified
scikit-learn/scikit-learn 23 22 95.7% Verified
sphinx-doc/sphinx 16 14 87.5% Verified
sympy/sympy 77 65 84.4% Verified
Consolidated Total (12 Repositories) 300 269 89.7% Physically Verified

Structural Context Compaction

In repository-scale benchmarks like SWE-bench Lite, large language models are overwhelmed by raw context windows exceeding 500,000 lines across hundreds of files. AtacamaODR compacts repository structures directly on-device into compact interface outlines. This preserves complete functional contract definitions while discarding internal implementation noise, allowing the 14B on-device model to evaluate repository architecture in sub-4-second horizons.

Economics of On-Device Localization

Frontier AI agents typically spend $0.50 to $2.00 in cloud prompt tokens simply reading repository directory structures to locate where an issue exists. By resolving 89.7% of fault localization targets on-device at $0.00 cost, AtacamaODR acts as an autonomous pre-flight filter, reserving expensive frontier cloud escalation solely for complex bounded code patch generation.

Verification Methodology & Ground Truth

This evaluation was executed against the official canonical SWE-bench Lite split (300 tasks). For each task, the problem statement was provided along with the compacted repository interface spine. The model predicted the target modification path, which was programmatically evaluated against the ground-truth commit diff from the official SWE-bench repository. Zero synthetic mock scripts or synthetic rate simulators were utilized.