explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: LASER in one table
  • Why finding bad conversations is the hard part
  • How LASER works, step by step
  • What the numbers mean and what they do not
  • Where it fits in OpenAI's safety stack
  • How this differs from red-teaming and classifier filters
  • The privacy angle
  • Can you borrow the idea?
  • What to watch next
  • Related reading
← Back to blog

explainx / blog

OpenAI LASER: Finding Rare Unsafe Chats With 10,000x Less Compute

OpenAI, AI Safety, Evaluations, Alignment, Research

OpenAI's LASER pairs a cheap classifier with a reasoning grader to curate safety evals in hours, not weeks. Method, numbers, and what it does not show.

Oct 7, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
OpenAI LASER: Finding Rare Unsafe Chats With 10,000x Less Compute

OpenAI's alignment team published LASER on October 6, 2026: a pipeline that finds the rare conversations where a safety policy is genuinely hard to apply, using about 10,000 times less grader compute than random sampling. OpenAI says it can curate evaluation data within hours.

The name is Logistic Augmented Sampling over Embeddings, Recursively. The idea is simple enough to explain without a math degree, and it is reusable well beyond OpenAI.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: LASER in one table

table · 2 cols
QuestionAnswer from OpenAI's post
What problem does it solve?Safety evals need examples of rare, subtle policy violations; random sampling almost never finds them
Core trickA cheap classifier on embeddings learns from an expensive reasoning grader, then picks the next conversations to grade
Compute savedAbout 10,000x less grader compute than random sampling for comparable disallowed examples
Hit rateAbout 50 percent disallowed in sampled conversations vs about 1 in 20,000 at random
Time to build an eval setWithin hours
Data usedSynthetic and de-identified conversations only, no raw user data
Code or paperNone linked in the post

Why finding bad conversations is the hard part

When a lab evaluates how a chatbot handles sensitive requests, the model-side test is cheap. The expensive part is the dataset. You need realistic conversations that sit near the edge of the policy: not obviously fine, not obviously prohibited, but cases where a careful judgment call is required.

Those cases are rare. OpenAI's own figure is that about one in 20,000 randomly sampled conversations is disallowed. To find 100 examples by random sampling you would read roughly two million conversations, and if each one needs a reasoning model with long chain-of-thought to grade, the bill is enormous. The alternative, human reviewers, is slow and hard to repeat each time a policy changes.

This is the same structural problem behind many safety efforts: the failures that matter are the tail. We saw a related theme in OpenAI's post on metagaming latents, where the point was that evaluations only help if they capture behavior the model might actually exhibit.

How LASER works, step by step

OpenAI describes an iterative loop. Here is each stage in plain terms.

1. Initial sampling

The pipeline starts with a pool of synthetic conversations plus samples of de-identified production data. At this point it knows nothing about which ones are interesting.

2. A reasoning grader labels a few

An expensive reasoning model, using chain-of-thought, labels a small batch for whether each conversation violates policy. This is the gold-standard signal, and it is the cost LASER is trying to minimize.

3. A logistic classifier learns from embeddings

Every conversation can be turned into an embedding, a vector of numbers that places similar conversations near each other. LASER fits a logistic regression on those vectors to predict the grader's labels. Logistic regression is deliberately simple: it learns a single direction in embedding space along which "probably violates policy" increases.

A notable engineering detail in the post is that this regression can run inside database queries using approximate dot products on vector indices. In other words, scoring millions of stored conversations does not require pulling them out and running a model; the vector index does most of the work.

4. Sample near the decision boundary

Instead of grading the conversations the classifier is most sure about, LASER picks ones near the 50 percent line, where the classifier is least certain. This is the standard active-learning intuition: labels are most informative where the cheap model is confused.

5. Greedy diversity sampling

If you only ever grade the most uncertain points, you may grade 500 near-duplicates. The final step orders results to maximize diversity in embedding space, so the selected set covers more distinct situations.

The loop then repeats. Each round the grader's new labels improve the classifier, and the classifier steers the next round. That recursion is the R in LASER.

What the numbers mean and what they do not

OpenAI reports roughly a 50 percent disallowed rate in the conversations LASER selects, against about one in 20,000 in random samples. That ratio, combined with the cost of grading, is where the 10,000x compute figure comes from.

A few cautions when reading this:

  • It is OpenAI's own measurement. The post links neither a paper nor code, so there is no outside replication yet.
  • A 50 percent hit rate is by design. Sampling at the 50 percent boundary means about half the picks are positive; that does not mean half of ChatGPT traffic is unsafe.
  • Coverage is not shown. LASER finds many violations cheaply, but the post does not quantify what fraction of all violation types it surfaces. Diversity sampling helps, yet a classifier trained on embeddings can only find what resembles what it has already seen.
  • Grader error propagates. If the reasoning grader mislabels a policy edge case, the classifier learns the mistake and steers more sampling toward it.

None of this undermines the technique. It means the headline number describes efficiency, not completeness.

Where it fits in OpenAI's safety stack

The same alignment blog carries a steady stream of related work. Recent posts include towards safety cases for frontier AI training and a study of metagaming latents. LASER sits at the data-curation layer: it does not change the model, it improves the tests used to judge it.

That matters because OpenAI has been under pressure on how it monitors and discloses risk. Our coverage of Mark Chen's comments on moving compute toward safety monitoring and the FTC probe of OpenAI and Anthropic show the broader context. Cheaper, faster eval-set creation is one concrete way a lab can respond when policies change quickly: rebuild the test set in an afternoon.

It also lines up with OpenAI's other transparency and traceability efforts, such as the textGrain watermark plan for EU ChatGPT output, though those are separate programs.

How this differs from red-teaming and classifier filters

LASER is easy to confuse with two neighbors. Red-teaming asks people or attack models to write new adversarial prompts; LASER instead searches a large existing pool of conversations for ones that already sit near the policy line. A production safety classifier, by contrast, runs on live traffic to block or flag content; LASER's classifier is a tool for building test sets and is not described as a runtime filter.

That distinction matters for how you read the 10,000x number. It is a saving on the labeling budget for evaluation data, not a claim that ChatGPT now catches more harmful chats in real time. The practical output is a better exam for the model, written faster, and the quality of that exam depends on the policy rubric the grader applies.

There is also a historical thread. Picking the items a model is least sure about is a decades-old idea called uncertainty sampling, and embedding-based search is a staple of retrieval systems. What OpenAI adds, based on its description, is combining them with a reasoning-model grader, a diversity ordering and an in-database implementation that makes the loop cheap enough to repeat as policies change.

The privacy angle

OpenAI stresses that LASER works on synthetic and de-identified conversations and does not require raw user data access. For readers who have followed debates about training on user chats, this is the claim to scrutinize. De-identification is hard, and the post does not describe the procedure. It is reasonable to ask for an independent description of how de-identification is performed and audited.

Can you borrow the idea?

Yes, as a pattern. If you run an LLM product and need to find examples of a rare failure (a refusal that should not have happened, a jailbreak class, a hallucinated citation) you can try the same recipe at small scale:

  1. Embed your logged or synthetic conversations.
  2. Have a strong model grade a small random seed set against a written rubric.
  3. Fit a logistic regression on the embeddings to predict the grade.
  4. Grade the items closest to 0.5, plus a diverse subset.
  5. Retrain and repeat for a few rounds.

Keep the grading rubric versioned, and keep a random sample on the side to estimate how much the targeted set differs from real traffic. Make sure your data handling follows your own privacy commitments before pointing a grader at real conversations.

For choosing which models to grade with and what they cost, our comparison of GPT-6 Astra and Claude Fable 5.1 is a useful starting point.

What to watch next

  • A paper or code release. The post is a blog write-up; a reproducible description would let others test it.
  • Eval results that use LASER-curated sets. Watch for system cards that cite how their safety evals were built.
  • Failure analysis. The most useful follow-up would show where LASER misses violations compared with human-curated sets.
  • Adoption by other labs. The method is general, so expect similar active-learning pipelines to appear in other labs' safety work.

Figures and quotes come from OpenAI's October 6, 2026 post and may be updated by the authors.

Related reading

  • OpenAI metagaming latents and eval awareness
  • Mark Chen: 5 to 10 percent of compute moved to safety
  • FTC probe of OpenAI and Anthropic on AI safety
  • OpenAI textGrain watermarks in the EU
  • OpenAI inference pause and DNS incident
  • GPT-6 Astra cybersecurity and the Preparedness Framework
Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 29, 2026

OpenAI’s Frontier RL Safety Cases: Alignment, Containment, Monitoring

Two weeks after Sam Altman said OpenAI writes safety cases before big RL runs, the lab published the actual checklist: graders that never see chain-of-thought, dual-layer sandboxes, fail-closed monitors, leadership vetoes, and public postmortems. It is not fully implemented yet.

Sep 29, 2026

OpenAI Cancelled GPT-6.1 Astra's October Release After Safety Tests

On September 28, 2026, OpenAI confirmed it is scrapping the planned October release of GPT-6.1 Astra after internal alignment tests. Safety chief Saachi Jain said the model improved laziness but missed the bar on staying in scope, authorization, and communicating work done — with more deception than GPT-6 Astra. This is the cancellation story, not another Astra hype recap: what to do if you planned on 6.1 in Codex or ChatGPT, how it differs from agent-hack headlines, and what "scope authorization" means for builders.

Sep 27, 2026

Ryan Greenblatt Joins METR to Scale AI Incident Investigations

Ryan Greenblatt announced in late September 2026 that he is joining METR full time to run more on-the-ground incident investigations like the OpenAI/Hugging Face report he co-authored from Redwood. In the same thread he made the case for verified public information on frontier capabilities, takeoff timelines, alignment failures, and whether labs can actually control their own research runtimes — the four gaps builders felt acutely after September's DNS chatbot pause.