Your gen AI model just gave a wrong, outdated, or hallucinated answer. What do you do? For most teams the reflex is "fine-tune it on our data." That is usually the wrong first move — it is the most expensive, slowest option, and it often does not even fix the problem.
This guide is a decision tree for choosing the cheapest fix that actually works. It is also one of the highest-value topics on the Google Cloud Generative AI Leader certification, where "fine-tuning by default" is an explicit trap.
First, diagnose the failure
Different failures need different fixes. Ask what actually went wrong:
- Wrong format or tone (right facts, bad structure) → a prompting problem.
- Outdated or missing facts (the model does not know your data or recent events) → a grounding / RAG problem.
- Consistent style or domain behavior you cannot get from prompting → a fine-tuning candidate.
- High-stakes correctness (clinical, legal, financial) → always add human-in-the-loop, regardless of the above.
The five tools, cheapest to most expensive
1. Prompt engineering — start here
Reshaping the input is the fastest, cheapest lever. Clearer instructions, role prompting, few-shot examples of the desired format, chain-of-thought for reasoning, or prompt chaining for multi-step tasks. If the model has the knowledge but formats it badly, prompting alone often fixes it. (See our prompting techniques guide.)
2. Grounding — anchor to authoritative data
Grounding connects output to trusted data so the model stops inventing facts. Google frames three sources:
- First-party data (your own documents, databases),
- Third-party data (licensed or partner sources),
- World / public data (e.g., Grounding with Google Search).
Grounding is the conceptual category; RAG is one way to implement it.
3. RAG — retrieval-augmented generation
RAG retrieves relevant documents at query time and injects them into the prompt, so answers are based on current source material without retraining the model. On Google Cloud this shows up as prebuilt RAG with Agent Search and RAG APIs. RAG is the workhorse for "make the model answer from our knowledge base." (For the retrieval mechanics, see our embeddings and vector search guide.)
Critical caveat: grounding and RAG reduce hallucinations — they do not eliminate them. The model can still misread or over-extrapolate from retrieved context.
4. Fine-tuning — change behavior, not currency
Fine-tuning adjusts the model on curated examples to shift its style, tone, or domain behavior. It is the right tool when prompting cannot get you consistent behavior — for example, a very specific output structure across thousands of calls. It is the wrong tool for keeping facts current: a fine-tuned model still has a knowledge cutoff, and re-tuning for every data change is costly. It is the most expensive and slowest option, so reach for it last.
5. Human-in-the-loop — the safety net
HITL keeps a person in the decision path. Two things people forget: it is not just final QA — HITL also supports continuous monitoring (catching drift, bias, and edge cases over time), and it is mandatory for high-stakes use cases no matter how good your grounding is.
The decision tree
- Is the failure about format, tone, or reasoning steps? → Prompt engineering.
- Is the model missing your data or recent facts? → Grounding / RAG.
- Do you need consistent domain behavior that prompting cannot deliver? → Fine-tuning (after 1 and 2).
- Is the output high-stakes? → Add human-in-the-loop and continuous monitoring on top of whichever fix you chose.
Notice the order: prompting and grounding are cheaper and faster than fine-tuning, and they solve the most common failures. Fine-tuning is powerful but narrow.
The responsible-AI and SAIF angle
Choosing a fix is not only a cost question — it is a governance question.
- Grounding and RAG improve transparency and explainability: you can cite the source an answer came from, which matters for accountability.
- Fine-tuning raises data-quality, privacy, and bias questions: what went into the tuning set, was it anonymized or pseudonymized, and could it bake in bias?
- HITL and continuous monitoring are how you operationalize responsible AI over time — not a one-time check.
- Google's Secure AI Framework (SAIF) applies throughout the ML lifecycle. Remember its six elements are non-sequential — you apply them together, not as a checklist you finish before launch. (More in our certification guide.)
Common traps
- Fine-tuning as the default fix. It is expensive, slow, and does not keep facts current. Try prompting and grounding first.
- "Grounding eliminates hallucination." It reduces, not eliminates. Keep HITL for high-stakes work.
- Confusing grounding and RAG. Grounding is the goal; RAG is one technique to achieve it.
- Treating HITL as final QA only. It is also part of continuous monitoring.
- Ignoring the governance implications of each fix — sources, privacy, and bias differ across them.
A worked diagnosis: an assistant gives the wrong refund policy
Imagine a support assistant that says customers have thirty days to request a refund, while the current policy allows fourteen. First inspect the request and the evidence the model received. If no policy was supplied, this is a missing-information problem. Adding a style instruction cannot create the missing policy, and fine-tuning yesterday's rules would leave the same maintenance problem tomorrow.
Now suppose retrieval supplied the current policy but also an older document. The failure may be document selection, conflicting authority, or unclear version metadata. Identify the authoritative source and retire or clearly label superseded content. Ask the assistant to cite the policy section it uses. This makes a wrong answer easier to diagnose without assuming every failure comes from model reasoning.
If retrieval supplied only the correct fourteen-day policy and the answer still says thirty, inspect how the context was assembled. Was the source truncated? Did an instruction encourage a generic answer? Did the user ask about a different product? Change one variable at a time so you can attribute any improvement to the actual intervention.
Separate retrieval quality from answer quality
Create an evaluation set with two expectations per case: which evidence should be retrieved, and what the answer should say. A question can fail because the right document was never found, or because the model misunderstood a correct document. Scoring only the final answer hides that distinction and makes debugging slower.
Include questions that the documents cannot answer. A grounded assistant should sometimes say it lacks enough information. If every case has a known answer, your evaluation cannot show whether the assistant invents a policy when evidence is absent. Add ambiguous questions that require clarification, such as a refund request without a purchase date.
When might fine-tuning actually be a reasonable next step?
Consider fine-tuning after you have a stable task definition and examples of the behavior you want. A recurring transformation with consistent output conventions may benefit from it. Keep factual source retrieval separate when facts change frequently. Fine-tuning and RAG can coexist because they address different parts of the system.
Do not assume a universal cost ranking. A small task-specific model, repeated long prompts, retrieval infrastructure, and human review all have different operating costs. The earlier order is a practical starting heuristic, not a law that fine-tuning must always cost more. Compare the full workflow under your expected volume before making an architecture decision.
What should the human reviewer actually review?
Give a reviewer the input, proposed answer, and supporting evidence. A person asked only to approve fluent text cannot efficiently catch a missing source or a misread exception. For the refund example, show the purchase date and the policy clause alongside the proposed action. Route disagreements to a named owner rather than letting the assistant repeatedly rephrase its answer.
Log the category of correction: missing evidence, wrong evidence, misunderstood evidence, formatting, or unauthorized action. Those categories guide the next fix. If most corrections involve outdated documents, improve content freshness. If they involve schema violations, improve output constraints. If they involve ambiguous policy, ask the policy owner to clarify the rule before training a model on conflicting examples.
A useful system improves from these observations without treating every disagreement as a new training example. Some disagreements expose product policy gaps. Others expose access-control requirements. Keep those decisions with the responsible owner instead of asking the model to silently choose a policy on the organization's behalf.
Related on explainx.ai
- HyDE and "don't classify, hallucinate" — a related technique that leans on an LLM's own priors instead of grounding, and where that trade-off breaks down
- Google Cloud Generative AI Leader certification guide — the full exam breakdown
- Certification study guide — domains, scenarios, task map
- Learning pathway — articles mapped to exam domains
- Zero-shot, few-shot, and chain-of-thought prompting — prompt engineering foundations
- Embeddings and vector search — how RAG retrieval works
- What is an embedding? Examples — interactive similarity demo
- Top 10 open & closed embedding models — model shortlist for RAG
- What is bias in AI — responsible AI depth
- AI skills every developer needs in 2026 — where RAG and fine-tuning fit in the full developer roadmap
Written for gen AI literacy and Generative AI Leader exam prep. Product names and framework details are summarized from Google Cloud's public messaging; verify current specifics on Google Cloud. explainx.ai is not affiliated with Google Cloud.
Update — October 3, 2026: Related: Adaption's Invent API for zero-seed synthetic data and Qwen censorship audit.
