explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • What is in the TermGrade release?
  • How were the tasks built and checked?
  • What did six models score?
  • What is the "band-central" rule?
  • What does the training recipe look like?
  • What did the training actually gain?
  • What can you find in the trajectories?
  • Do difficulty labels transfer between models?
  • How do you use it, and what are the limits?
  • Bottom line
  • Related reading
← Back to blog

explainx / blog

TermGrade: 1,004 Open RL Environments for Terminal Agents, Explained

TermGrade, Reinforcement Learning, Terminal Agents, Datasets, Open Source AI

Part of AI Research

ai& released TermGrade: 1,004 graded terminal environments, 36,144 trajectories and an RL recipe that lifted gemma-4-31B-it 3.1 points on Terminal-Bench 2.1.

Oct 9, 2026·11 min read·Yash Thakker
add explainx.ai
go deep
TermGrade: 1,004 Open RL Environments for Terminal Agents, Explained

Training a terminal agent needs thousands of tasks that a machine can run and grade. Most of those task sets are private. On October 8, 2026, ai& released one in the open: TermGrade, with 1,004 executable Linux tasks, six models' pass rates for every task, all 36,144 attempts behind those rates, and the RL run that used them.

Yagiz Calik, who posts as Weyaxi, announced it on X: "high-quality RL environments for terminal agents, fully open source, along with the full story of how we built and graded them, and the RL recipe we trained with." This post is based on the ai& blog post and the Hugging Face dataset card.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

What is in the TermGrade release?

The blog post lists four Hugging Face repositories under ai-and, plus a checkpoint.

table · 2 cols
RepositoryWhat it holds
ai-and/termgrade-environments1,004 executable Linux tasks: container, instruction and pytest verifier, with a measured pass rate for six models. Apache-2.0
ai-and/termgrade-trajectories36,144 graded attempts: every command, raw terminal frames, per-test verdicts, reasoning where exposed. 21,910 pass and 14,234 fail
ai-and/termgrade-bandcentral-gemma4-31b-bashThe 151 training and 120 validation tasks, as exact parquet splits
ai-and/termgrade-gemma4-31b-lora and -mergedThe resulting adapter and the merged weights

Headline numbers from the post:

table · 2 cols
MetricValue
Executable environments1,004
Graded trajectories36k (36,144 counting trials)
Models graded6 (6,024 task-by-model cells)
Computeabout 20k hours and 4.5B tokens of agent rollout (the post's summary tile)
Verifier size16.2 named assertions per task on average

How were the tasks built and checked?

Reward hacking: a ball slipping under an obstacle, the failure that execution-verified RL tasks guard against

The pipeline has four stages, per the post.

  1. Task generation. DeepSeek-V4-Pro wrote the whole task from a domain and difficulty spec, with a repair loop. A task is four things: environment/, tests/, instruction.md and solution/.
  2. Filtering by execution. The team built each image offline and ran the task's own reference solution against its own verifier. They also audited for unsatisfiable tests that no implementation could pass. Of 66,165 candidate task directories, 1,004 survived (1.5 percent).
  3. Release grading. Six models attempted every task under the Terminus-2 scaffold, through the Harbor harness. Eight attempts (k=8) for most models and two for the costliest two, under a 12-hour scoring budget.
  4. Training-band grading. A separate measurement, using the base model alone in a bash tool-calling scaffold, selected the training data.

Grading is built to resist cheating. The dataset card says tests are not in the container during the attempt. They are copied in from a clean host-side copy at scoring time, so the agent cannot read what it is graded on. Grading runs offline (--network none) in a freshly built container. The card says the team checked for overlap with Terminal-Bench 2.1 and for near-duplicates and found none.

You can run one task yourself. The card gives a quick start: build the base image once (docker build -t dsbase-venv312:local base_image/), build a task, run its reference solution inside the container, copy the tests in and read /logs/verifier/reward.txt. Note the card's warning: the image has no test files, so you must copy the suite in yourself.

What did six models score?

The release grading, on the Terminus-2 scaffold, gave these pass@1 numbers:

table · 3 cols
Modelkpass@1
google/gemma-4-31B-it80.4986
google/gemma-4-31B-it (thinking)80.5774
Qwen/Qwen3.6-27B80.6320
moonshotai/Kimi-K380.6770
deepseek-ai/DeepSeek-V4-Pro20.6853
zai-org/GLM-5.220.6858

The post says Kimi-K3, DeepSeek-V4-Pro and GLM-5.2 are statistically tied (paired differences of +0.05, -0.87 and -0.82 percentage points, with p values of 0.97, 0.42 and 0.51). Qwen3.6-27B is separated from all three at p below 0.001. The DeepSeek numbers refer to the build ai& graded, not the 0813 release that shipped during their window.

The cost detail is the interesting part. A tie on score hides a large gap in time. Kimi-K3's median winning trial took 5.6 minutes, against 92.9 and 106.9 minutes for the other two tied models, and about 5,600 output tokens per win against roughly 100,000 and 73,000. Eight attempts per task for Kimi-K3 cost 1,929 trial-hours. Two attempts each for GLM-5.2 and DeepSeek-V4-Pro cost 11,648. For more on how those models do on this style of benchmark, see our coverage of DeepSeek V4 Pro on Terminal-Bench and GLM-5.3 on Terminal-Bench 4.

What is the "band-central" rule?

This is the core idea, and it is simple.

In GRPO, the gradient a prompt contributes scales with the variance of its rollout rewards. With a binary reward, that variance is p(1-p), where p is the model's pass rate on the task. It peaks at p = 0.50 and is zero at 0 and 1. A task the model always solves or never solves gives every rollout the same reward, so the update has nothing to compare. A task solved about half the time gives successes to reinforce and failures to compare them against.

A token budget jar, a reminder that RL rollouts for terminal agents cost billions of tokens

So the rule is: select the tasks whose measured base pass rate is closest to 0.50, with no hand-curation. The steps in the post:

  • Several thousand task-gradings against the base model.
  • 555 tasks in the 0.25 to 0.75 range.
  • The 151 closest to 0.50. That is 118 from the TermGrade corpus plus 33 from SETA, a synthetic terminal-environment set from CAMEL-AI.
  • Of the 151, 137 carried gradient once training began. Fourteen never produced a passing rollout.

Two more findings from the post matter for practitioners:

  • Volume did not help. A 707-task loosely in-band corpus scored 46.6 against 46.1 for the 151-task selection. The difference was not distinguishable from zero (t = 0.80). Nearly five times the data bought no measurable advantage.
  • The band moves. The mean pass rate on the prompts that carried gradient was 0.36 at training time, not 0.50. The post's lesson: a difficulty label is one measurement of one model at one moment.

What does the training recipe look like?

The team calls it GRPO++: GRPO with DAPO modifications for long-horizon agentic RL and Dr.GRPO's advantage handling. From the post:

table · 2 cols
SettingChoice
KL termNone, in the reward or as a loss
ClippingClip-Higher, asymmetric PPO clipping at 0.2 and 0.28
Loss aggregationToken-mean, so long trajectories weigh in proportion to length
AdvantagesMean-centered only, no standard-deviation normalization
SamplingDynamic: drop prompts whose 16 rollouts all pass or all fail, then resample (up to 10 batches per step)
Batch16 prompts, 16 rollouts each, temperature 0.6, fully on-policy
Base model and adaptergemma-4-31B-it, LoRA rank 32, learning rate 1e-5, binary reward
Length of run5 updates, about one epoch
HardwareOne 8xH200 node, about 7 hours (4 training, 3 validation)

Training ran in a bash tool-calling scaffold and was evaluated under Terminus-2, so the cross-scaffold transfer is part of the result. For a wider view of RL environment tooling, see our guide to Hugging Face Hub RL environments, OpenEnv and verifiers and the Agent Lightning v1 guide.

What did the training actually gain?

The released checkpoint scored 46.1 on Terminal-Bench 2.1, 3.1 points above the base 43.0. The base figure comes from Artificial Analysis. The 46.1 is the mean of seven evaluation runs with three attempts per task. The team followed Artificial Analysis' method: 89 tasks, Terminus-2, pass@1 over three repeats and a two-hour per-task timeout, with Daytona through Harbor as the sandbox instead of E2B.

The more honest number is the repeat run. Five trainings of the identical recipe, with no fixed seed, averaged 45.1, which is +2.1 with a standard error of 0.6. All five beat the base (the weakest by 0.4 points, the strongest by 3.3). The gain is distinguishable from zero (t = 3.8, p about 0.025). The authors say the released +3.1 is one run, and +2.1 is the better estimate of the method.

They also say validation, a 120-task held-out set in a different scaffold, moved with the benchmark in four of five runs and against it in one, so Terminal-Bench is the measurement of record. And the post notes the results rest on a small number of runs.

What can you find in the trajectories?

Most released agent trajectory sets include only successes. TermGrade keeps every attempt, and the post shows why the failures matter. In one example, gemma-4-31B-it builds a fixed-point calculator in twelve turns, passes 44 of 45 assertions and fails test_empty_input. Its reward is 0.0, "indistinguishable from a model that never started." The post says 6,544 trials (18 percent) failed exactly one assertion while passing eight or more. The team tried per-test partial credit and found it worse than binary reward on this corpus, and calls it the most interesting unexplored direction.

Each trial ships in two forms: an SFT-ready messages column and the full Terminus-2 episode with every keystroke and terminal frame, plus per-turn token metrics and named assertion verdicts. Observations the authors list:

  • gemma-4-31B-it with thinking ends a third of its trials in two turns or fewer, writing a script and declaring completion without running it.
  • Score and cost move independently. pass@1 spans 1.4x across the models while median tokens per win spans 34x.
  • GLM-5.2 emits turns with no valid tool call, and those turns run longer than its successful ones.
  • More turns is not progress: within a task, winning trials used fewer turns than failing ones for Qwen and GLM.

Do difficulty labels transfer between models?

No, and this is the most useful warning in the release. The authors measured the "solved about half the time" set for all six models. The Jaccard overlap between any two different models' sets runs 0.11 to 0.21, and not one of the 1,004 tasks is band-central for all six. The highest overlap, 0.29, was gemma-4-31B-it against itself with thinking on. The post concludes that difficulty belongs to the combination of task, model and scaffold, so you grade against the model you plan to train. That is why every trajectory ships: so you can redo the measurement.

How do you use it, and what are the limits?

  1. Pull ai-and/termgrade-environments and read labels.jsonl (per task and model: pass_rate, band, n_counted, k).
  2. Filter to split == "unseen" (802 tasks) if you benchmark. 202 tasks were used in ai&'s own train or validation splits (113 train151 and 89 heldout120).
  3. Regrade the tasks against your own model and scaffold, then pick the tasks near 0.50.
  4. Use the 21,910 passing trajectories as an SFT set if you want, since the messages column feeds a chat template unmodified and the license is Apache-2.0.
  5. Verify a task with its reference solution before you trust it (oracle_verified flags the ones where that could not be confirmed).

Limits to keep in mind:

  • Small number of runs. The authors flag it themselves.
  • One base model. The recipe was tested on gemma-4-31B-it only.
  • Model-generated tasks. DeepSeek-V4-Pro wrote the tasks, so they carry that model's habits. Execution filtering checks that a task is solvable, not that it is representative of your workload.
  • Scaffold dependence. Terminus-2 and the bash scaffold give different pass rates, which the post states directly.
  • Secret scanner noise. Hugging Face's scanner reports Lob API keys on the dataset. The card says they are false positives (pytest function names that start with test_).
  • We did not rerun it. explainx.ai did not reproduce the training or the benchmark.
  • No developer thread. We searched Hacker News for TermGrade and found no discussion, so we report none.

Bottom line

TermGrade's value is less the gemma gain than the data hygiene. Every task was proven solvable by running it. Every pass rate points to raw trials. The authors admit that difficulty is local to a model and a scaffold, and they ship the raw attempts so you can redo the measurement. If you train terminal agents, regrade the corpus against your own model and keep the 802 unseen tasks clean for evaluation.

Related reading

  • Agent harness engineering on Terminal-Bench (LangChain)
  • Terminal-Bench 2.0 benchmark guide
  • Hugging Face Hub RL environments, OpenEnv and verifiers
  • Agent Lightning v1 and agentic RL
  • Fermisense GRPO 9B catalog
  • DeepSeek V4 Pro 0813 on Terminal-Bench
  • What is an agent harness?

Sources

  • ai&: TermGrade, 1,004 graded terminal environments for RL (October 8, 2026)
  • Hugging Face: ai-and/termgrade-environments, and the linked termgrade-trajectories repository
  • Weyaxi's announcement on X
  • HuggingNews summary

Figures are from the ai& post and dataset card as read on October 9, 2026. Terminal-Bench numbers depend on the scaffold, sandbox and repeat count. Verify them before you cite them.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 17, 2026

Xiaomi MiMo-V2.6: Livestreaming a Trillion-Parameter RL Training Run

Fuli Luo's Xiaomi MiMo team announced on September 17, 2026 that MiMo-V2.6 is mid-run on a large-scale reinforcement learning training pass, and is livestreaming it publicly — with plans to open-source scaling details on compute, environments, and grading over the coming weeks.

Jul 27, 2026

Intelligence Ownership: $500 Fine-Tune Beats Frontier

A July 27 Fermisense case study claims a ~$500, 3.5-day GRPO run on a 9B open model beat five frontier configs on scored catalog integrity — and crushed unit economics. explainx.ai extracts the playbook and the skepticism.

Oct 9, 2026

Cactus Whistle: A 16.9 MB Speech-to-Text Model That Runs on a CPU

On October 2, 2026, Cactus Compute released Whistle, a 16.9 MB Apache 2.0 speech recognition model that runs on one CPU engine beside the Needle tool-calling model. It beats Whisper base on several benchmarks and loses on others. Here is the full picture.