Sakana AI announced on October 9, 2026 that its Japanese-focused language model, Sakana Namazu, has been adopted by Evidence Finder. Evidence Finder is a literature search tool for physicians from Aillis Inc. (アイリス株式会社). The announcement carries one headline number: a Namazu-based evaluation model scored 96.4% on Japan's 120th national physician examination.
The post is a customer story, not a new model release. It is useful for two reasons. It shows a Japanese-tuned model working as the answer layer in a real medical-search product. And it shows, by the company's own wording, why an exam score alone should not be read as proof that an AI is ready for clinical work.
An earlier request to explain "Namazu / AILLIS" confused the names. AILLIS is the customer, Aillis Inc. Namazu is Sakana's model.
TL;DR: what was announced
| Question | Answer |
|---|---|
| Who is involved? | Sakana AI (model: Namazu) and Aillis Inc. (product: Evidence Finder) |
| What happened? | Evidence Finder adopted Sakana Namazu to write comparative, cited answers for doctors |
| Headline number | 96.4% on Japan's 120th national physician exam (February 2026), per Aillis |
| Caveat from Sakana | An exam score shows one facet of knowledge and reasoning, not clinical usefulness |
| What does the model do? | Compares and synthesizes papers that Aillis's algorithms select, then writes the answer |
| What checks the sources? | A Verify function that confirms cited papers exist |
| Is there a new model? | No. The post describes the current Namazu |
| Pricing or API details? | None in the post |
What is Evidence Finder?
Evidence Finder is a service from Aillis, a Japanese medical AI company. According to Sakana's post, it "searches literature from databases such as PubMed in response to a doctor's question and generates answers while naming the sources." It has a Verify function that automatically checks whether the papers cited by the AI actually exist. The design goal is to let doctors "reach the primary literature quickly."
The need is plain. Sakana's post says clinicians must find, read and compare findings from papers and guidelines that update every day, but doctors whose main job is treating patients have limited time for literature research. A tool that finds papers and writes a cited summary saves time, but only if the citations are real. Hallucinated references are the known failure of general chatbots. The Verify step targets exactly that. For background on why models invent sources and how to catch them, see explainx.ai's guide on why AI models hallucinate.
Japanese trade press reported earlier that Aillis agreed to use Sakana's model in Evidence Finder on September 14, 2026, according to a Yakuji Nippo summary. Sakana's own post is dated October 9.
Where does Namazu fit in the pipeline?
The post splits the work between two parties.
- Aillis algorithms choose relevant papers for the doctor's question. This is the search and verification layer that Aillis built.
- Sakana Namazu compares and integrates those papers and writes the answer.
This is a retrieval-augmented design. The retrieval layer finds the evidence. The language model writes and organizes it. If you build similar systems, the split is worth copying. It limits what the model can claim to what the retrieved papers say. explainx.ai covers the tradeoffs in RAG versus agentic RAG and in grounding with RAG or fine-tuning.
Sakana says the adoption reflects its Japanese answer ability and the medical performance described next. In Sakana's words, the model "naturally" fits the workflow of getting doctors to trusted primary literature.
What does 96.4% on the physician exam mean?

The post says the Namazu-based evaluation model for Evidence Finder scored a 96.4% correct rate on the 120th National Medical Licensing Examination, held in February 2026. A Japanese trade summary of Aillis's own announcement gives the raw score as 482 out of 500 points.
Aillis says this is the highest score among domestic foundation models designated by the Ministry of Economy, Trade and Industry's GENIAC program, among those with publicly confirmable results. The footnote explains the method: as of September 1, 2026, Aillis compared the physician exam results from the 118th exam onward that are available in each company's published materials. That is a comparison by Aillis, among published results, and not an independent leaderboard.
Sakana adds its own warning. The post says: "The national exam result shows only one facet of knowledge and reasoning ability, and does not directly show usefulness in clinical practice." The decision to adopt Namazu also rested on a qualitative review of real answers, which found practical usefulness.
That is the right way to read it. A licensing exam tests recall and reasoning on multiple-choice questions. Real clinical questions are open, messy and tied to a patient. The exam score is a screening signal. It is not an outcome study.
What is Namazu, and how is it adapted to Japanese?
Sakana describes Namazu as a Japanese-specification LLM. It starts from a strong open model. The post names the problem that often follows when you train a high-performing open model further on Japanese: "catastrophic forgetting," where the model loses existing knowledge and reasoning, or becomes over-optimized for narrow tasks.
Sakana says Namazu addresses this with its own post-training technique. It keeps the base model's reasoning, knowledge and coding ability on major benchmarks while strengthening adaptation to Japanese language, Japanese cultural and business context, and neutral responses. The current Namazu maintains the base model's performance in mathematical reasoning, general knowledge and reasoning, and coding. It beats the base model on Japanese instruction following, translation and evaluations about Japan-specific context.
The post states the goal: not "just a model that is good at Japanese," but a way to make "the capabilities of the world's best open models usable in Japan." Sakana also says it is building a base where inference completes inside Japan, including data handling and operations, so Japanese companies and organizations can choose it more easily.
The post does not name the base model or give benchmark tables. explainx.ai covered the earlier product surfaces in Sakana Chat adds Fugu, new Namazu and code execution. That August post noted that no new public benchmark numbers came with the update. The 96.4% exam score is the first concrete domain number attached to a Namazu deployment that we have seen.
Why does this matter beyond Japan?
There are three lessons for people who build with AI.
Domain wrappers beat raw models. The value in Evidence Finder is the pipeline: selected papers, a writer model, a source check. Namazu is a part of it. A smaller or regional model, well tuned for language and tone, can be the right writer when the retrieval is strong.
Language adaptation without losing skill is a real engineering problem. Catastrophic forgetting is why many local-language fine-tunes feel weaker at reasoning. Sakana claims its post-training avoids that. The claim is not independently benchmarked in this post, but the target is correct, and the approach connects to the guide on what fine-tuning is and when it helps.
Sovereignty has a medical use. Japan wants models that run domestically with clear data handling. Healthcare is a place where that preference is strong. This links to the broader debate in What Is Sovereign AI? and to Europe's parallel path with models such as Aleph Alpha Kolibri.
Sakana is also widening where Namazu applies. The post lists human resources, marketing, customer support and internal research as areas that need Japanese systems and business-practice knowledge. It says Namazu is strongest at "accurately understanding Japanese instructions and materials, conducting consistent work from information retrieval through organization and explanation." Interested companies can consult via a form.
What are the limits and open questions?
- The score is vendor-reported. Aillis ran the comparison. Sakana repeats it with attribution. No independent test is cited.
- It is an evaluation model. The post says "the evaluation model of Evidence Finder using Sakana Namazu" scored 96.4%. The production model may differ. The post does not say.
- No clinical outcome data. There is no study of whether doctors using the tool make better decisions or save a measured amount of time.
- No error analysis. The post does not report how often Verify catches a fake citation, or how often the summary misreads a paper.
- Base model is unnamed. You cannot reproduce or compare the adaptation without it.
- Medical advice caution. Evidence Finder is described for physicians. It is not a patient-facing diagnosis tool. People should not use any AI search summary in place of medical advice. For a patient-side view of AI and health records, see ChatGPT Health and your medical records.
- Source language. We read the Japanese original of Sakana's post.
How to evaluate a similar tool yourself

If you are choosing or building an evidence-search assistant, five checks give you more than any exam score.
- Citation existence. Take 50 answers and check each citation by hand. Count fakes.
- Citation support. For real citations, check that the paper says what the answer claims.
- Recency. Ask about a guideline updated in the last six months. See if the tool finds the new version.
- Refusal behavior. Ask something with no good evidence. A good tool says so.
- Language quality. Ask in the user's real language, with the real jargon, and have an expert read the output.
What people are asking
Is Namazu a medical model? No. Namazu is a general Japanese-specification model. Evidence Finder is the medical product. The exam score belongs to the Namazu-based evaluation model.
Does Namazu replace doctors' search? It is designed to shorten the path to primary papers, with the doctor still reading the source.
Is this the same as Sakana Fugu? No. Fugu is Sakana's multi-model orchestrator, covered in Sakana Fugu: one API to orchestrate the others. Namazu is the Japanese-first model.
Related reading
- Sakana Chat adds Fugu, new Namazu and in-chat code execution
- Sakana Fugu: one model API to orchestrate all the others
- Sakana Fugu Max and Ultra v2
- What Is Sovereign AI?
- Why AI models hallucinate and how to catch it
- RAG versus agentic RAG
- What is fine-tuning?
Primary sources: Sakana AI: Sakana Namazu adopted by Evidence Finder (Japanese); Aillis Inc.; Yakuji Nippo summary of the September 14 agreement (Japanese).
Figures come from Sakana AI's October 9, 2026 post and Aillis's published comparison as of September 1, 2026. They are company claims and may be updated.
