explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

  1. Home
  2. /
  3. Dictionary
  4. /
  5. Abliteration
Safety & Alignmentaka Refusal Removalaka Obliteration

Abliteration

Abliteration is a weight-level technique that finds the direction in a model's activations responsible for refusals and removes it, producing a permanently uncensored model without retraining.

Ask Melo about this← all terms

Abliteration computes a "refusal direction" by averaging the activation difference between harmful and harmless prompts across a model's layers, then orthogonally projects that direction out of the weights so it no longer influences outputs. Unlike fine-tuning, it needs no training data or GPU training run — it is an inference-time computation applied once, permanently, to the model's weights. Tools like Heretic automate the process; commercial vendors including OrcaRouter and Abliteration.ai have applied it to models like Qwen3.8-27B and GLM-5.3, either as downloadable weights or as a hosted API marketed to red teams and cybersecurity testers.

Related terms

JailbreakModel RefusalAI SafetyCapability ControlSandboxingMechanistic Interpretability