Adaption Labs has opened API access to Invent, a system that generates post-training datasets from a text description — no seed examples required. Its technical report, "Invent a Dataset: Measuring dataset generation abilities with zero seed data," claims about 17 percent higher quality and up to 37 percent more diversity at 20,000 rows than frontier models used as data generators, with 0.0 percent exact-duplicate prompts.
If those numbers hold on your task, they change what you can do without a data team. They are also vendor-run results, so the right move is to understand what was measured and test it on your own workload.
TL;DR — what people are asking
| Question | Answer |
|---|---|
| What is it? | A hosted API that creates SFT or preference datasets from a prompt |
| Who built it? | Adaption Labs (paper authors include Sara Hooker) |
| Seed data needed? | No — "zero seed data" is the point |
| Reported quality gain? | About 17% relative over the strongest baselines |
| Reported diversity gain? | 19% at 2K rows; 37% at 20K rows |
| Duplicates? | 0.0% exact-duplicate prompts at all scales; baselines 6.2–19.2% |
| Output formats? | JSONL, JSON, CSV or Parquet; instruction pairs or preference pairs |
| Limits? | Text only; no tool-call or agentic trajectories |
| Price? | Not in the report; credit-based per a third-party summary |
| Biggest caveat? | Results are from the vendor's own benchmark |
The problem it targets
Fine-tuning a model starts with data, and teams often have none. You know the behavior you want — answer legal questions in a particular style, handle banking support tickets, write Hindi news summaries — but there is no labeled dataset. The usual workaround is to prompt a frontier model to generate examples. That works at small scale and degrades at large scale: outputs converge, prompts repeat, and diversity collapses.
Adaption's paper frames this as the zero-data regime and builds a benchmark around it: given only a natural-language description, how good is the dataset a system can generate? For the fundamentals of why data quality drives fine-tuning results, see our guide to what fine-tuning is and the decision framework in grounding, RAG and fine-tuning.
How the API works
Based on the product description, the workflow is asynchronous:
- Submit a request through the
datasets.inventcall with a prompt of up to 10,000 characters, a domain code (broad likemedicalor narrow likemedical.symptoms_diagnosis), a target dataset size, and an output type. - Receive a dataset ID.
- Poll the job status until it completes.
- Download the result as JSONL, JSON, CSV or Parquet, structured as instruction pairs for supervised fine-tuning or preference pairs for preference optimization.
Billing is credit-based, with maximum row counts set by plan tier, and language expansion costs extra based on a sample_rate setting. There is no documented self-hosted option.
What the benchmark measured
The paper evaluates eight task types:
| Task | Why it is a hard test |
|---|---|
| African QA | Underrepresented in typical training data |
| Legal QA | Precision matters; hallucination is costly |
| Serbian car ads | Narrow language and genre |
| Customer support, banking | Domain vocabulary and policy |
| Hindi news | Non-English, formal register |
| Customer support, general | Broad intent coverage |
| Medical QA | Safety-sensitive |
| Hotel reviews | Style and sentiment variation |
Diversity is scored on both semantics and wording: an embedding-based score called DCScore, plus n-gram diversity, compression ratio and exact-duplicate rate. Quality is scored on prompt quality, completion quality and relevance on a 0–10 scale, and by ranking models fine-tuned on the data using an LLM as judge.
The comparison set includes proprietary reference generators — Claude Opus 5, GPT-5.6 Sol and Gemini 3.1 Pro — and open-weight generators that are permitted for training, GLM-5.3 and DeepSeek V4 Pro.
The headline results, in context
- Quality: roughly 17 percent relative gains over the strongest baselines.
- Diversity: 19 percent relative gain at 2,000 samples, widening to 37 percent at 20,000.
- Duplicates: 0.0 percent exact-duplicate prompts across scales, versus 6.2 to 19.2 percent for baselines, rising with dataset size.
- Under constraints: a reported 58 percent advantage over the next best, GLM-5.3, when restricted to training-permissible generators.
- Post-training: models fine-tuned on Invent data were ranked first on 54 percent of prompts for Llama-3.3-70B and 41 percent for Gemma-4-31B-it.
The diversity pattern is the most plausible and useful part. It matches the known failure mode: the gap is near parity at about 200 samples and grows with scale, which is exactly where naive prompting collapses. One third-party write-up notes that Claude produced less varied data at 20,000 than at 2,000 examples.
Why you should still be skeptical
These results come from the authors, on a benchmark they designed, scored partly by an LLM judge. That does not make them wrong, but three things deserve a check:
- Benchmark fit. A system built to win a zero-seed-data benchmark may be tuned to its metrics. Your task may weigh things the benchmark does not.
- Judge bias. LLM-as-judge scoring has known biases; the paper's post-training rankings inherit them.
- Baseline strength. A frontier model with a well-engineered generation pipeline — seeded personas, deduplication, rejection sampling — can close part of the gap. The comparison may use simpler prompting as the baseline.
What it cannot do
The authors are direct about limits, which is a good sign:
- No tool-call traces or agentic trajectories. If you are training an agent to use tools, this is not that.
- Text only. Multimodal extension is future work.
- SFT and alignment data only. It does not produce harness datasets.
- An open question. The authors note it is unclear whether this kind of data is best absorbed through weight updates at all, versus other approaches.
Generated data also needs validation. Coverage of the product warns that synthetic data requires checking for factual errors and unsafe content before production use. In a domain like medicine or law, that is not optional.
How to evaluate it in an afternoon
You can learn more from a small, honest test than from the report.
- Pick one narrow behavior you already know how to judge, such as a support-ticket classifier or a style-constrained summarizer.
- Generate a pilot set of 1,000 to 2,000 rows with Invent, and an equivalent set with your current method (a frontier model plus your best prompt).
- Measure duplicates and diversity yourself. Exact-duplicate rate and a simple embedding-spread check take minutes.
- Hand-review 100 rows from each. Look for factual errors, repeated templates and off-topic rows.
- Fine-tune a small model on each and test both on a held-out set you wrote yourself.
- Compare cost. Include credits, review time and re-generation.
- Scale only if it wins. The diversity advantage is claimed to grow with size, so a second test at 10,000 rows is worth the spend if the pilot is promising.
If you are building a post-training pipeline more broadly, our overview of open-source RL-as-a-service stacks shows where data generation fits next to training and evaluation, and Prime Intellect's Prime Inference launch covers the serving side of the same loop. For the distillation angle — and the terms-of-service questions around using one model's outputs to train another — read what AI distillation is.
A note on licensing and terms
The paper describes the endpoint as permitting downstream commercial use, and its comparison focuses on generators that are permissible for training. That matters because many frontier APIs restrict using their outputs to train competing models. If you generate data with a frontier model today, check its terms. If you use Invent, confirm that the license for the generated data matches your intended use and keep a copy of the terms with the dataset.
What people are asking
Does this replace human-labeled data?
No. It reduces the cold-start problem when you have no data. For high-stakes domains you still need expert review, and for evaluation sets you should use human-written data, not generated data.
Will models fine-tuned on synthetic data collapse?
Model collapse is a risk when models train repeatedly on their own outputs. Diversity is the main protection, which is why the paper emphasizes it. Mix synthetic data with real data when you can.
Is a prompt of 10,000 characters enough to describe a behavior?
It is a generous limit for a specification. Treat the prompt as a spec: include format, tone, edge cases and examples of what to avoid.
Can I use it for non-English tasks?
The benchmark includes Hindi news, Serbian car ads and African QA, which suggests multilingual intent. Language expansion is a billed option, so check cost per language.
How does it differ from asking GPT or Claude for examples?
Per the report, the difference shows up at scale: fewer duplicates and a diversity gap that widens past a few thousand rows. At small scale the report shows near parity.
Honest limitations
- All performance numbers are from Adaption's own technical report.
- Pricing is not published in the report; third-party summaries may be out of date.
- I have not run Invent myself.
- Data-handling terms were not reviewed in full.
Bottom line
Invent offers a hosted way to create SFT and preference data from a description alone. The claimed edge — fewer duplicates and a diversity gap that grows with dataset size — targets a real failure of prompt-based generation. Test it on a narrow task with your own held-out evaluation, review the rows by hand, and check the license before scaling.
Related on explainx.ai
- What is fine-tuning? Complete guide
- Grounding, RAG or fine-tuning: a decision guide
- Open-source RL-as-a-service post-training stacks
- What is AI distillation?
- Prime Inference launch
- AI evals for engineers and PMs
- How to read AI benchmarks
Sources: Invent a Dataset technical report (arXiv) · AlphaSignal summary, October 1, 2026
Results and product details reflect Adaption Labs' report and third-party summaries as of October 3, 2026.
