Ask ten AI-103 candidates to define semantic, vector, and hybrid search and you will get ten overlapping, half-right answers. This is the single most confusable topic on the AI-103 exam, and it shows up across both the generative/agentic and information extraction domains because it underpins RAG grounding. Here is the clear version.
The three approaches, in one paragraph each
Keyword (lexical / BM25) search — the baseline
Classic keyword search ranks documents by term overlap, weighted by term frequency and rarity (the BM25 algorithm). It is exact and fast, and it excels at precise tokens — product IDs, error codes, names, SKUs. Its weakness: it does not understand meaning. Search "car" and it will not match "automobile" unless you have synonyms configured. It is the foundation the other approaches build on.
Vector search — meaning by embeddings
Vector search converts both your query and your documents into embeddings (high-dimensional numeric vectors) and ranks by similarity (typically cosine distance) in that space. Because embeddings capture meaning, "car" and "automobile" land near each other, so vector search handles paraphrases and synonyms well. Its weakness is the mirror image of keyword search: it can miss exact tokens — a specific invoice number or a rare product code may not be well represented in embedding space. For the foundations, see our embeddings & vector search guide.
Semantic search — reranking with language understanding
Semantic search (in Azure AI Search, the semantic ranker) is a reranking step applied on top of an initial result set. It uses a language model to re-score and reorder candidate results by how well they actually answer the query, and it can surface captions and answers. Crucially, semantic ranking does not replace retrieval — it improves the ordering of whatever candidates the retrieval step returned. Think of it as a quality pass, not a retrieval strategy by itself.
Hybrid search — combine keyword + vector (+ optional semantic rerank)
Hybrid search runs both keyword (BM25) and vector retrieval and fuses their results (Azure AI Search uses Reciprocal Rank Fusion to merge the two ranked lists). You get keyword precision (exact IDs, rare terms) and meaning-based recall (paraphrases, synonyms) in one query. Add the semantic ranker on top and you have the strongest default for RAG grounding.
Under the hood: what actually happens
| Approach | Retrieval signal | Handles synonyms? | Handles exact IDs/codes? | Extra step |
|---|---|---|---|---|
| Keyword (BM25) | Term overlap + frequency | No (without synonyms) | Yes | — |
| Vector | Embedding similarity | Yes | Weak | Embed query + docs |
| Semantic | Reranks existing candidates | Improves ordering | Improves ordering | LM reranking pass |
| Hybrid | BM25 and vector, fused | Yes | Yes | Rank fusion (RRF) |
The key mental model: keyword and vector are retrieval strategies; semantic is a reranking layer; hybrid is the combination of the two retrieval strategies. Semantic ranking can be layered on any of them, most powerfully on hybrid.
When to use which
- Exact-match heavy (codes, IDs, legal citations, SKUs) → keyword, or hybrid to also catch meaning.
- Natural-language, paraphrase-heavy questions → vector, or hybrid to also catch exact tokens.
- You need the best ordering of results → add semantic reranking.
- General RAG grounding for an agent or app → hybrid + semantic reranking is the reliable default.
The exam tell: a scenario describing "combining keyword precision with meaning-based relevance" is hybrid — not "semantic alone." Candidates who pick semantic there are conflating the reranking layer with the retrieval combination. Similarly, if a scenario says "users search with synonyms and paraphrases and miss exact-keyword documents," the fix is usually to move from keyword-only to hybrid (or add vector), not to crank up chunk size.
How Azure AI Search implements all three
Azure AI Search is the documented retrieval engine for AI-103, and it supports the whole stack in one service:
- Index your content with both a searchable text field (for BM25) and a vector field holding embeddings.
- Issue a keyword, vector, or hybrid query. Hybrid queries run both and fuse results with Reciprocal Rank Fusion.
- Optionally enable the semantic ranker to rerank the fused results and return semantic captions/answers.
- Enrich during indexing with built-in skills (native) or custom skills (a hosted function/API you call in the skillset) — for chunking, embedding generation, OCR, and more.
This ties directly into RAG ingestion: chunk documents, generate embeddings, attach grounding metadata, and index — then ground your Foundry app or agent on hybrid + semantic results. For the ingestion side, see our RAG pipeline design guide.
Common mistakes to avoid on the exam
- Calling semantic a retrieval strategy. It is a reranking step on top of retrieval.
- Picking semantic when the scenario says "keyword precision + meaning." That is hybrid.
- Assuming vector search handles exact IDs well. It often does not — that is where hybrid earns its keep.
- Confusing built-in vs custom skills in the enrichment pipeline — custom skills require a hosted function/API.
- Reaching for fine-tuning when a retrieval upgrade (keyword → hybrid, or adding semantic rerank) is the real fix.
Walk through an exact-code query and a paraphrase query
Imagine a support collection with two articles. One explains how to replace the battery in device QX-17. Another describes a charging problem using phrases such as "will not hold power" and "runs out quickly." A customer asks, "Why does my QX-17 die after charging?" The query contains both a precise identifier and an imprecise description of the symptom.
Keyword retrieval can preserve the identifier. Vector retrieval can connect the symptom to different wording in the document. Hybrid retrieval gives both signals an opportunity to contribute candidates. Semantic reranking can then improve the order of the retrieved passages. This example illustrates the roles; it does not establish that one configuration will always outperform another on your collection.
Now ask only for a warranty claim with a specific invoice number. A result about similar warranty processes may be topically relevant but still be the wrong record. Exact identity and user access should be requirements, not suggestions to the embedding model. Keep identifiers in suitable fields and evaluate those queries separately from broad natural-language questions.
The Azure hybrid-search overview describes a single request combining full-text and vector queries, with optional semantic ranking. Use that implementation reference when translating the conceptual example into a real index and API request.
Evaluate retrieval before blaming the answer model
Build a small question set containing exact codes, paraphrases, ambiguous requests, and questions the collection cannot answer. Label the source passage that should support each answer. Run the same questions through keyword, vector, and hybrid configurations while holding the indexed documents constant.
Inspect whether the relevant passage enters the candidate set before reranking. A ranker cannot rescue a document that never reached it. If the correct passage is absent, examine ingestion, chunk boundaries, embedding compatibility, filtering, and candidate limits. If it appears but ranks poorly, compare the reranking stage and the text fields it receives.
Keep retrieval and final-answer evaluation separate. An answer can be wrong despite excellent retrieval, and fluent output can disguise poor retrieval. Record the returned passages, their source identifiers, and the final answer so you can identify which stage changed after a configuration update.
Do not compare scores from different ranking stages directly
A keyword score, vector similarity, fused rank score, and reranker score describe different operations. Their numerical scales need not be interchangeable. A larger number from one query type does not, by itself, prove that its answer is better than another query type's answer.
Compare results against relevance judgments instead. For a help-center application, ask whether the returned passage answers the specific user's question and whether the top results cover the required facts. For a record lookup, ask whether the exact authorized record was retrieved. Those success criteria make configuration changes interpretable.
Finally, preserve access filters throughout the experiment. A configuration that improves recall by including documents the caller should not read is not an improvement for that caller. Search quality includes selecting the correct permitted evidence, rather than maximizing apparent similarity at any cost.
Check the embedding contract when changing models
Document embeddings and query embeddings must use the compatible representation expected by the index. If you switch the embedding model, check dimensions and regenerate the affected indexed vectors as required. A query returning results does not prove that mismatched embeddings produce meaningful similarity. Preserve a small reference query set and compare the retrieved passages after the change, rather than accepting a successful HTTP response as the quality check.
Bottom line
Vector = meaning by embeddings. Semantic = a language-model reranking pass. Hybrid = keyword and vector fused, and the best RAG default when you add semantic reranking. Azure AI Search does all of it in one engine. Nail this distinction and you neutralize the exam's favorite retrieval trap. Keep going with the full AI-103 exam guide, the certification study guide, and the learning pathway.
For embedding intuition and model choice, see what is an embedding (with examples) and the top 10 open & closed embedding models shortlist.
Behavior described reflects Azure AI Search as of early 2026; verify current features on Microsoft Learn. explainx.ai is not affiliated with Microsoft.
