AtacamaODR Apple Silicon
BENCHMARK AUDIT: TRACK 04 • CANONICAL BLIND HUMANEVAL (164 PROBLEMS) • AUTHOR: GANESH NALLASIVAM
Download Track 04 Report (.PDF)

Canonical Blind HumanEval Benchmark

Full 164-problem blind execution benchmark evaluating Qwen2.5-Coder-14B-Instruct (4-bit) running on Apple Silicon via AtacamaODR (Port 8088). Every test case was executed in an isolated sandboxed subprocess with strict 5.0-second timeouts against official OpenAI canonical assertions. Zero synthetic mock data. Zero prompt overfitting regexes. 100% physically executed code.

Pass@1 Accuracy
86.0%
141 of 164 passed.
Total Problems Evaluated
164 / 164
100% completed suite.
Average Turnaround
13.8 s
Total runtime: 37.7 min.
Execution Failures
23
19 assertion, 4 syntax errors.
Industry Standard Python Coding Accuracy (HumanEval Pass@1)
Grounded Empirical Comparison
Model Architecture Parameters / Quant Host Environment Pass@1 Accuracy Status
Qwen2.5-Coder + AtacamaODR 14B (4-bit MLX) Apple Silicon (24GB UMA) 86.0% (141/164) Physically Verified
Claude 3.5 Sonnet (Frontier Reference) Unknown (Cloud) Anthropic Cloud API 93.7% Published Lab Metric
GPT-4o (Frontier Reference) Unknown (Cloud) OpenAI Cloud API 90.2% Published Lab Metric
Llama 3.1 70B Instruct 70B (16-bit / 8-bit) Enterprise Cloud Host 80.5% Published Lab Metric
Qwen2.5-Coder-7B-Instruct 7B (4-bit) Local Apple Silicon 79.9% Published Model Card
Audit Methodology & Error Breakdown

Empirical Execution Telemetry

All 164 problems were piped via HTTP to http://127.0.0.1:8088/v1/chat/completions, parsed for Python function blocks, concatenated with canonical test cases, and executed in temporary isolated files via python3 subprocesses with 5.0-second timeouts.

Passed Unit Tests:
141 Problems
Exit Code 0 (Pass@1)
Logical / Assertion Errors:
19 Problems
Canonical assertion mismatch
Syntax / Parsing Errors:
4 Problems
Unterminated strings / syntax
Dataset: OpenAI Canonical HumanEval.jsonl (164 tasks) Artifact: humaneval_execution_results.json