explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • What the list actually says
  • TL;DR — what people are actually asking
  • The Hermes omission is the real story
  • Is B-tier for Claude Code and Codex defensible?
  • Is Oh My Pi actually S-tier material?
  • The tiers nobody argued about — and why that's suspicious too
  • No methodology, no rank — and someone asked directly
  • What to do instead of copying someone else's tier list
  • Related on explainx.ai
← Back to blog

explainx / blog

A Viral Agent Harness Tier List Put Claude Code in B — Does It Hold Up?

Agent Harness, Claude Code, Codex, Cursor, AI Agents

A viral tier list ranked Oh My Pi S-tier and Claude Code/Codex B-tier — with Hermes missing entirely. Here's what holds up and what doesn't.

Sep 22, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
A Viral Agent Harness Tier List Put Claude Code in B — Does It Hold Up?

A tier-list screenshot with no methodology attached is not evidence — it's a conversation starter, and this one started a big one. On September 21, 2026, an X user posting as Sayo shared an image ranking agent harnesses from S to F, captioned simply "ai agent/harness tiers." It reached nearly 80,000 views within hours, and the replies disputed almost every placement in it — starting with the harness that isn't on the list at all.

What the list actually says

table · 2 cols
TierHarnesses
SOh My Pi
AFactory Droid, Pi, Zed, fx
BCodex, OpenCode, Cursor, Amp, Devin, Claude Code, Grok Build
CWarp
DRoo Code, Cline
FAntigravity, GitHub Copilot

No scoring rubric, task set, or usage disclosure accompanied the image — just the ranking itself, dropped without further comment beyond a follow-up post noting a dislike of "the over-use of emoji in default setup" on one entrant.

TL;DR — what people are actually asking

table · 2 cols
QuestionDirect answer
What topped the list?Oh My Pi, alone in S-tier
What's the biggest omission?Hermes — not ranked anywhere, called out repeatedly in replies
What's the most disputed placement?Claude Code and Codex in B-tier, alongside Cursor and Grok Build
Is there a scoring methodology?No — one commenter asked directly and got no answer
Is OMP actually S-tier material?Disputed — one reply called it "slop" vs. plain pi
Should you pick a harness from this list?No — treat it as one workflow's opinion, build your own eval
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

The Hermes omission is the real story

The single most repeated reaction across the replies wasn't about any specific tier placement — it was that Hermes, Nous Research's open agentic harness, doesn't appear on the list at all. At least four separate commenters asked variations of "where Hermes?" independently, without prompting each other, which is a stronger signal of a genuine gap than a single complaint would be. One reply went further, saying they'd switched their entire workflow to Hermes for coding and calling both the Hermes omission and Claude Code/Codex's B-tier placement "beyond dumb" in the same breath.

That reaction matters because Hermes has a real, documented technical profile — a remote-VPS-and-Telegram-CLI-oriented open harness with its own design philosophy distinct from every entrant that did make the list. Omitting an actively-used harness with its own committed user base isn't a minor oversight in a list this specific; it's evidence the ranking reflects one person's own tool exposure rather than a survey of the actual landscape.

Is B-tier for Claude Code and Codex defensible?

This is where the pushback gets more substantive than "you forgot my favorite tool." One reply called grouping Codex and Claude Code into B-tier, below OpenCode in the same tier and below a themed config wrapper in S-tier, "beyond dumb" — not because B-tier is a bad grade in isolation, but because it flattens a real distinction. explainx.ai's own comparison of the leading harnesses treats Claude Code and Codex as full-featured, terminal-first, deeply extensible systems — hooks, skills, subagent orchestration, MCP tool integration — that sit in a genuinely different design category from Cursor, an IDE-integrated product built around a different interaction model entirely. Putting all three in one bucket answers "did I use this today" more than it answers "which harness has the deepest extensibility surface" or "which harness handles a 200-turn autonomous session most reliably" — different questions that would plausibly produce different rankings.

Devin, also placed in B-tier, is a fully autonomous agent product rather than a developer-driven harness in the same sense as the terminal tools around it — another example of the list treating meaningfully different product categories as directly comparable on one axis.

Is Oh My Pi actually S-tier material?

The single sharpest technical disagreement in the thread was about the top slot itself. Oh My Pi (OMP) — a themed configuration layer built on top of Mario Zechner's minimal pi harness, following the same naming convention as Oh My Zsh for shell configs — got one blunt reply: "OMP is slop unfortunately stock pi with bare minimum for your workflow is much better." Sayo's own response defended it while separately admitting a dislike for OMP's default emoji usage, which is a strange combination of endorsing the S-tier ranking while criticizing one of its defaults in the same reply.

The underlying disagreement is a real one in harness design more broadly: does a themed, opinionated configuration layer on top of a minimal harness represent genuine added value, or does it just add surface area over a tool that was already good specifically because it shipped with a small, unopinionated core? Pi's own design philosophy — no baked-in MCP, no default sub-agents, extend only what you need — is arguably in tension with a themed wrapper being ranked a full tier above the base tool it's built on.

The tiers nobody argued about — and why that's suspicious too

It's worth noting what didn't draw pushback, because silence in a reply thread this contentious is its own signal. Warp sat alone in C-tier without comment. Roo Code and Cline — both mature, widely-used open-source VS Code extension agents — landed in D-tier with no defense mounted for them, despite both having real adoption in the wild. Antigravity and GitHub Copilot in F-tier, the list's harshest grade, also went unchallenged in the visible replies, even though Copilot in particular has one of the largest install bases of any tool on the entire list by a wide margin.

That's a pattern worth sitting with: the placements that generated real argument were the ones near the top (S and A tier) and the one most people recognized as clearly wrong (Hermes's absence), while the bottom two tiers passed without a single visible defender. Either the community broadly agrees Antigravity and Copilot deserve F-tier — a real possibility, since Copilot's agentic capabilities have historically trailed purpose-built coding-agent harnesses even as its install base dwarfs them — or the people who'd disagree simply didn't see the post or didn't bother replying. A tier list's uncontested placements aren't automatically correct; they're just placements nobody happened to push back on in this specific thread, which is a different and weaker form of validation than it looks like at a glance.

No methodology, no rank — and someone asked directly

Perhaps the cleanest single critique in the replies came from a commenter who simply asked Sayo: "this is based on what? What is your rank?" That question — never answered in the visible thread — is the actual problem with any viral tier list like this one. A ranking with no disclosed task set, no usage volume per harness, and no scoring rubric is indistinguishable from a personal preference dressed up as a verdict. Another reply put it more bluntly: "No way you have used all of these harnesses intimately enough to put them in tiers."

That's not a reason to dismiss the list — Sayo's bio describes them as a "software factory manager," a role that plausibly involves real exposure to several of these tools — but it is a reason to treat it as one practitioner's dashboard, not a benchmark result. It doesn't disclose what tasks were run, on what codebases, at what volume, or over what time period, which are exactly the variables that would make one harness outperform another for a specific reader's actual workflow.

What to do instead of copying someone else's tier list

  • Build your own ranking from your own repository and tasks. A harness that excels at long autonomous refactors may lose badly on quick interactive fixes, and vice versa — the same discipline covered in how to read AI benchmarks applies directly to harness comparisons, not just model benchmarks.
  • Separate "harness design philosophy" from "which one I happened to open today." Terminal-first extensible harnesses (Claude Code, Codex, OpenCode), IDE-integrated products (Cursor, Zed), fully autonomous agents (Devin), and minimal ownable cores (pi) are different categories solving different problems — a single ordinal ranking across all of them answers a narrower question than it appears to.
  • Check whether a harness you rely on is simply missing before trusting a list's floor as much as its ceiling. Hermes's total absence here is the clearest example of why a tier list's silence about a tool tells you more about the list's author than about the tool.
  • Use a structured comparison when you actually need to decide. explainx.ai's top 10 harnesses guide breaks the same landscape down by license, extensibility, and workflow fit rather than collapsing it into one axis.

Related on explainx.ai

  • Top 10 Closed-Source and Open-Source Agent Harnesses (2026)
  • Pi Agent Harness: Mario Zechner's Minimal Coding Agent You Can Own
  • Hermes Agent: Nous Research's Remote VPS + Telegram CLI Guide
  • What Is Harness Engineering for AI Agents?
  • How to Read AI Benchmarks
  • What Is an Agent Harness? Complete Guide
  • AgentRun: How Grep.ai Turns Agents Into Cheap, Auditable Workflows

This post reflects a viral X tier-list image and its replies as of September 21-22, 2026. Tier placements, reply quotes, and view counts are as they appeared on X at time of writing; the list's author disclosed no scoring methodology, and this post does not attempt to independently verify or replicate the ranking. Follow @explainx_ai for updates.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 4, 2026

Armature Study: What Claude Code, Codex, and Cursor Actually Pick

Armature, a startup that sells "growth services to dev tools," measured 16,893 coding-agent sessions to see which tools Claude Code, Codex, and Cursor actually pick — not just mention. The findings are genuinely useful (repo language flips winners, mentions don't equal picks) and the source is a genuine conflict of interest. Here's both, held at once.

Sep 18, 2026

Top 10 Harness Engineering Concepts Every AI Builder Should Know

Claude Code, pi, and Hermes look different on the surface but solve the same ten underlying problems. These are the concepts that separate a demo that falls over after ten turns from an agent you can trust to run unattended — each with a concrete example and a way to build it yourself.

Sep 18, 2026

What Is Harness Engineering? The Layer That Turns a Model Into an Agent

Claude Code, pi, and Hermes all call the same model APIs. What separates a working coding agent from a demo that falls over after ten turns is everything wrapped around the model: the agent loop, the tool contracts, the context and memory system, and the recovery logic that keeps a session alive across failures. That layer now has a name — harness engineering. Here's what it actually covers.