All posts
ai-agentsbenchmarksharnessproduction

Prime Agent Cleared the Human Expert Baseline by 0.1 Points. Same Model.

95.5% on ARC-AGI-3 against a 95.4% human expert baseline — running Opus 5, the model everyone already has. The harness rewrites itself mid-session, and it's the harness that moved.

NeuroX AI · August 8, 2026

Prime Agent hit #1 on GitHub Trending on August 7 and took +2,293 stars in a single day. The number worth reading is further down the page: 95.5% on ARC-AGI-3, against a human expert baseline of 95.4% — running Opus 5, the same model you can already call. Across three runs it held [95.0, 95.2, 95.5], and finished all 183/183 levels at Best@3.

The model didn't change. The harness did.

Two ideas carry it. The Recursive Language Model keeps context as variables in a persistent IPython kernel — file reads, shell, subagents, and compaction all happen as code, so state survives the turn instead of being re-narrated into the prompt every time. The Continual Harness stores supplemental prompts, memories, and skill specs as durable state, and /refine updates them mid-session from evidence in the trajectory.

Then the discipline: /refine never touches the immutable base system prompt. Only the supplemental layer, session-local by default. A self-improving agent with a fixed floor — which is why it converges instead of drifting. Prime Intellect reports it beat each model's own native harness on the long-context suite while using fewer total tokens.

The README stays honest about the cost: the kernel is lifecycle isolation, "not a security sandbox." It executes model-generated Python with your permissions. The capability ships in a weekend. The containment is still yours to build.

See how we ship it →

Contact

Working on something similar?

Tell us about it — we reply within one business day.

Or skip the form — book a Calendly slot directly

We reply within one business day · NDA on request

admin@neuroxai.com · +91 70149 99768

Remote-first team across India · US · EU · HQ in Udaipur, India