Index
- Exam overview
- Plan and manage an Azure AI solution
- Generative AI and agentic solutions
- Computer vision
- Text analysis and speech
- Information extraction
- If the question says X, pick Y
- Common traps
- Self-check questions
- Resources
Exam overview
AI-103 is Foundry-first. More than half the score (55-65%) is planning/securing Foundry and building gen AI apps and agents, so most of the study time should go there.
- Pass mark is 700/1000.
- Questions are scenario-based, Python-flavored, and mostly on GA features (common preview features can appear).
- The shift from AI-102: stop thinking “which Cognitive Service?” and start thinking “which Foundry capability, model, or tool?”
- The old services now show up as Foundry Tools (Language, Speech, Translator, Vision, Content Understanding, Document Intelligence) inside a Foundry resource + project.
Exam domains
| Domain | Weight | What it covers |
|---|---|---|
| Plan and manage an Azure AI solution | 25-30% | Model/service choice, deployment types, quotas, security, responsible AI |
| Generative AI and agentic solutions | 30-35% | RAG, Foundry Agent Service, tools, multi-agent, evals, tracing |
| Computer vision | 10-15% | Image/video generation and editing, multimodal understanding, Content Understanding |
| Text analysis (incl. speech) | 10-15% | LLM extraction, Language, Translator, Speech |
| Information extraction | 10-15% | Azure AI Search, skillsets, Content Understanding for documents |
If you are coming from AWS
| Azure / Foundry | Closest AWS mental model |
|---|---|
| Foundry resource + project | Bedrock account setup + a workspace boundary |
| Foundry Agent Service | Bedrock Agents / AgentCore Runtime |
| Foundry Tools catalog, MCP tool, OpenAPI tool | Action groups, AgentCore Gateway |
| Azure AI Search | OpenSearch / Kendra / Bedrock Knowledge Bases |
| Azure AI Content Safety, guardrails | Bedrock Guardrails |
| Content Understanding | Bedrock Data Automation |
| Foundry evaluations + tracing (App Insights) | Bedrock evaluations + AgentCore Observability |
Exam day tips
- Read the last sentence of the question first. It tells you what is being asked (cheapest, least effort, most secure).
- Flag the question and move on if it is taking more than 90 seconds.
- Case studies can’t be revisited once you leave them.
Plan and manage an Azure AI solution
Most questions here are “pick the right option under a constraint” (cost, latency, residency, security). Know the object model, the deployment types, and the security defaults cold.
Foundry object model
- Foundry resource (kind
AIServices): The top-level Azure resource. Holds model deployments, Foundry Tools, networking, keys/identity. - Foundry project: The working boundary inside the resource. Holds agents, connections, evaluations, traces, files.
- Apps connect with the project endpoint:
https://<resource>.services.ai.azure.com/api/projects/<project> - Use
AIProjectClient+DefaultAzureCredential
- Apps connect with the project endpoint:
- Connections: How a project reaches outside resources (Azure AI Search, Storage, Bing grounding, MCP servers, APIs).
- Credentials live in the connection, not in code or prompts.
Choosing a model
| Need | Pick |
|---|---|
| General chat, strong reasoning, tool calling | Flagship LLM (GPT-4.1 / GPT-5 class) |
| Hard multistep reasoning, math, planning | Reasoning model (o-series / GPT-5 reasoning) |
| Cheap, fast, edge or offline, simple tasks | Small language model (Phi family, mini/nano variants); Foundry Local for on-device |
| Images + text in one call | Multimodal model (GPT-4o / 4.1 class) |
| Vectors for search | Embedding model (text-embedding-3-small/large) |
| Generate or edit images | gpt-image-1 class, FLUX |
| Generate video | Sora class |
| Low-latency voice conversation | Realtime / audio models, or Speech + LLM |
| Prebuilt task (translate, OCR, PII, STT) | A Foundry Tool, not an LLM |
- The smallest model that meets quality wins cost questions.
- A Foundry Tool beats a prompt when the task is standard and deterministic output matters.
- Model names change often. Where a model is named, treat it as “the current model of that class”.
Deployment types
| Type | Use when |
|---|---|
| Global Standard | Default. Pay per token, highest quota, data may process in any Azure region |
| Data Zone Standard | Pay per token but processing must stay in the US or EU data zone |
| Standard (regional) | Processing must stay in one region |
| Provisioned (Global / Data Zone / Regional), PTUs | Predictable latency and throughput for steady high volume; reserved capacity |
| Global Batch / Data Zone Batch | Large async jobs, about 50% cheaper, results within 24 hours |
| Serverless API vs managed compute | Partner/open models: serverless = pay per token, managed compute = you pay for VMs |
Quotas, scale and cost
- Quota is TPM (tokens per minute) and RPM per model, per region, per subscription.
- Hitting the quota returns HTTP 429.
- Honor the
retry-afterheader - Use exponential backoff
- Honor the
- Scale beyond one deployment: more regions/deployments behind Azure API Management as an AI gateway.
- Load balancing
- Token-limit policy
- Token metrics
- Semantic caching
- Provisioned can spill over to Standard.
- Cost levers:
- Smaller model
- Batch
- Prompt caching
- Shorter prompts /
max_tokens - Semantic cache
- PTU reservations for steady load
Monitoring
- Models/agents:
- Azure Monitor metrics (tokens, latency, 429s)
- Application Insights tracing via OpenTelemetry
- Foundry observability dashboards
- Continuous evaluation on production traffic
- Search:
- Indexer execution history and errors
- Index size / document count
- Query latency
- Relevance testing
Security
- Keyless: Microsoft Entra ID + managed identity,
DefaultAzureCredential, disable local (key) auth. Keys only in Key Vault if you must. - RBAC: Least privilege.
- A user/app role to call models and agents (Azure AI User / Cognitive Services OpenAI User)
- A higher role to create deployments and manage the resource
- Grant the Foundry managed identity roles on the AI Search and Storage it reads
- Network: Private endpoints, disable public network access, VNet injection for agents.
- Agent setup:
- Basic = Microsoft-managed storage
- Standard = bring your own Cosmos DB (threads/conversations), Storage (files), AI Search (vectors)
- Pick Standard for compliance, residency, or private networking
- CI/CD: Azure Developer CLI (
azd) + Bicep/Terraform, GitHub Actions or Azure DevOps. Run evaluations as a pipeline gate before promotion.
Responsible AI
- Content filters / guardrails: Applied to input and output.
- Harm categories: hate, sexual, violence, self-harm
- Severity levels: safe, low, medium, high
- Add-ons:
- Prompt Shields = user jailbreaks + indirect/document attacks
- Groundedness detection
- Protected material (text and code)
- Custom blocklists
- PII detection
- Task adherence for agents
- Evaluators:
- Quality: groundedness, relevance, coherence, fluency, similarity, retrieval
- Safety: harmful content, indirect attack, protected material, code vulnerability
- Agent: intent resolution, tool call accuracy, task adherence
- AI Red Teaming Agent (PyRIT) for adversarial testing.
- Auditing: Trace logs, provenance metadata (C2PA content credentials on generated images), approval workflows.
- Agent governance:
- Tool allow-lists
- Per-tool auth
require_approval="always"on MCP tools- Human-in-the-loop for high-impact actions
- Agent identity in Entra
Generative AI and agentic solutions
The biggest domain. If you only master one thing: which agent tool solves which problem, and how grounding (RAG) is wired into agents.
RAG
- Ingest → chunk → embed → index (Azure AI Search)
- At query time, run hybrid search (keyword + vector) + semantic ranker
- Put the top chunks in the prompt
- Answer with citations
- Evaluate groundedness
- Agentic retrieval (Azure AI Search knowledge bases / Foundry IQ): Newer pattern where an LLM plans and runs subqueries for you. Use it for complex multi-part questions.
Agent building blocks
- Agent = model + instructions (role, goal, constraints) + tools.
- Kinds of agents:
- Prompt agent: declarative, no code hosting
- Workflow: multi-step orchestration, visual or YAML
- Hosted agent: your containerized code (e.g., Microsoft Agent Framework or LangGraph)
- Conversation state:
- New API uses conversations + Responses API
- Classic API used thread → message → run
- Classic agents are deprecated and retire March 31, 2027, but exam questions may still show thread/run code. Know both vocabularies.
- Memory:
- Short-term = the conversation
- Long-term = agent memory store (preview) or your own store (Cosmos DB)
- Trim or summarize long histories to fit context
- Tool schemas: Function tools are JSON Schema (name, description, parameters, required). Good descriptions drive correct tool selection.
Agent tools
| Scenario | Tool |
|---|---|
| Answer from a few uploaded files, zero infra | File Search (managed vector store) |
| Answer from an existing enterprise index | Azure AI Search tool (via project connection) |
| Current public info from the web | Grounding with Bing / Web search |
| Math, data analysis, charts, file transforms | Code Interpreter (sandboxed Python) |
| Call your own logic in your app process | Function calling (model returns a call; your code runs it and submits output) |
| Call an existing REST API with a spec | OpenAPI tool (anonymous, API key via connection, or managed identity) |
| Reuse a remote tool server | MCP tool (server_label, server_url, require_approval, connection for auth) |
| Low-code business workflow | Azure Logic Apps / Azure Functions |
| M365 docs or Fabric data | SharePoint tool / Fabric data agent |
| Multi-step web research report | Deep Research tool |
| Let a main agent delegate | Connected agents / A2A |
Multi-agent orchestration
- Connected agents: An orchestrator agent calls specialist agents as tools. Simple delegation, no custom code.
- Workflows / Microsoft Agent Framework patterns:
- Sequential
- Concurrent (fan-out/fan-in)
- Group chat
- Handoff
- Human-in-the-loop checkpoints
- Agent Framework is the successor to Semantic Kernel + AutoGen.
- A2A protocol is for agents across platforms; MCP is for tools.
- Safeguards for autonomy:
- Approval steps before irreversible actions
- Max iterations / turn limits
- Tool allow-lists
- Content filters on both input and output
- Task adherence checks
Code shape to recognize
from azure.identity import DefaultAzureCredential
from azure.ai.projects import AIProjectClient
project = AIProjectClient(endpoint=PROJECT_ENDPOINT,
credential=DefaultAzureCredential())
# new API: define agent (model + instructions + tools),
# then call it through the OpenAI-compatible Responses API
# classic API: create_agent -> threads.create -> messages.create -> runs.create_and_process
- When code uses an API key where Entra would work, the “most secure” answer swaps in
DefaultAzureCredential.
Evaluations
| Question | Evaluator |
|---|---|
| Fabrication / hallucination | Groundedness (is the answer supported by retrieved context?) |
| Did retrieval work? | Retrieval / relevance |
| Is it well written? | Coherence, fluency |
| Agents | Intent resolution, tool call accuracy, task adherence |
- Run with the
azure-ai-evaluationSDK or the Foundry portal, on a test dataset. - Compare runs, wire into CI/CD, enable continuous evaluation in production.
- Error analysis = read the traces of failed cases.
Optimize and operationalize
- Order of fixes: prompt engineering → RAG (knowledge gaps) → fine-tuning (style/format/behavior, or distill to a cheaper model)
- Fine-tuning types: SFT, DPO (preference), RFT (reasoning models)
- Parameters:
temperatureortop_p= change one, not both- Low temperature for extraction/factual, higher for creative
max_tokens, stop sequences, frequency/presence penaltiesreasoning_efforton reasoning models- Structured outputs (JSON schema) for reliable JSON
- Prompting: Clear system message, few-shot examples, delimiters, ask for citations, chain-of-thought for multistep tasks.
- Reflection / self-critique: generate → critique (same or judge model) → revise, with a stop condition. LLM-as-judge is how evaluators work.
- Observability: OpenTelemetry tracing to Application Insights. Spans per model call and tool call show token usage, latency breakdown, and safety signals.
- Hybrid orchestration:
- Route by task (small model for classification, big model for reasoning)
- Use a rules engine for deterministic compliance checks and the LLM for language
Computer vision
Three buckets: generate/edit media, understand media with multimodal models or Content Understanding, and keep it safe.
Generate and edit
- Images (gpt-image-1 class, FLUX):
- Text-to-image
- Image + reference images
- Edits with a mask (inpainting): the transparent area of the mask PNG is what gets regenerated
- Controls: size, quality, number of images, background (transparent), output format
- Video (Sora class):
- Text-to-video, image-to-video, and edit/remix an existing video
- It is an async job: create job → poll status → download
- Controls: resolution, duration, number of variants
- Generated media carries C2PA content credentials (provenance metadata) so it can be identified as AI-generated.
Understand images and video
- Multimodal chat: Pass images as
image_url(public URL or base64 data URI) in the message content.detail: low= cheap, fastdetail: high= fine detail, more tokens- Multiple images in one message for comparison
- Captions: Concise vs detailed is controlled by the prompt (and
max_tokens). - Alt text:
- Short and functional (about one sentence, no “image of”)
- Decorative images get empty alt
- Complex images (charts) get an extended description
- Visual Q&A: Instruct the model to answer only from what is visible and say when it can’t tell.
- Locate objects/regions: Content Understanding or Azure AI Vision object detection give bounding boxes. Plain LLM chat is weaker at precise coordinates.
Azure Content Understanding
A Foundry Tool where you define an analyzer with a field schema (or use a prebuilt one) for documents, images, audio, or video.
- Returns structured fields + markdown, with confidence and grounding.
- Video: segments, shot/scene detection, key frames, transcript, per-segment fields.
| Mode | Use when |
|---|---|
| Standard (single-task) | One file, one extraction |
| Pro | Multi-file, cross-document reasoning, can use reference data for validation; complex, multi-step extraction |
Responsible AI for images
- Content filters apply to image inputs and outputs (hate, sexual, violence, self-harm).
- Indirect prompt injection via text in images (a photo containing “ignore previous instructions”):
- Treat image text as untrusted data
- Extract it with OCR and scan with Prompt Shields
- Keep system instructions separate
- Limit which tools the agent can call
- Visual policy rules:
- Watermarks/provenance on generated media
- Detect prohibited symbols or brand misuse with custom analyzer fields or custom categories
- Block or route for human review
Text analysis and speech
The recurring question is “LLM prompt or Foundry Tool?”
- Pick the Tool for standard, repeatable, auditable tasks (PII redaction, sentiment at scale, document translation).
- Pick the LLM for flexible schemas, nuance, and tone.
Text analysis
- Structured JSON from an LLM: Use structured outputs (
response_formatwith a JSON schema, strict) instead of “please return JSON”. Low temperature for extraction. - Azure AI Language (Foundry Tool):
- NER
- PII detection and redaction
- Key phrases
- Sentiment + opinion mining
- Language detection
- Summarization
- Custom NER/classification
- Conversational language understanding (CLU)
- Question answering
- Safety and sensitive content: Azure AI Content Safety for harmful text; Language PII for personal data.
- Domain customization (compliance summaries, domain extraction):
- System prompt with glossary and rules
- Few-shot examples
- Output schema
- Custom NER or fine-tuning if prompts aren’t enough
Translation
| Need | Pick |
|---|---|
| Translate strings in real time, many target languages in one call | Azure Translator text translation |
| Translate whole files, keep formatting | Translator document translation (async, Blob Storage source/target, managed identity) |
| Company terminology | Custom Translator (train on parallel data) or glossary |
| Tone, style, context-aware rewrites | LLM-powered translation flow |
Speech
- Speech to text:
- Real-time (streaming)
- Fast transcription (synchronous, files)
- Batch (large volumes, async)
- Diarization for who-spoke-when
- Accuracy on domain terms:
- Phrase list first (no training, quick)
- Custom speech model (train with text and audio) when accents, noise, or vocabulary need more
- Text to speech:
- Neural voices
- SSML controls pronunciation, pauses, rate, pitch, speaking style
- Custom neural voice is limited access (approval needed)
- Speech as an agent modality:
- Voice Live API (low-latency speech-to-speech, can front a Foundry agent)
- Realtime audio models
- The classic STT → agent → TTS pipeline
- Handle barge-in and turn detection
- Reasoning over audio: Audio-input models, or Content Understanding audio analyzers (transcript + extracted fields).
- Speech translation: Speech service translation (speech to text/speech in target languages) or STT → LLM translate → TTS.
Information extraction
This is Azure AI Search plumbing plus Content Understanding for documents. Know the indexer pipeline order and the four query types.
Azure AI Search pipeline
- Data source: Blob, ADLS, SQL, Cosmos DB, SharePoint, etc.
- Indexer: Pulls on a schedule, change and deletion detection, field mappings. Check execution history for errors/warnings.
- Skillset (enrichment):
- Built-in skills:
- OCR
- Image Analysis
- Document Layout (layout-aware, markdown chunks)
- Text Split (chunking)
- Text Merge (put OCR text back into content)
- Entity recognition, key phrases, language detection
- Azure OpenAI Embedding
- Custom skill: Web API skill (usually an Azure Function) with a fixed JSON input/output contract
- Built-in skills:
- Index:
- Fields with attributes: key, searchable, filterable, sortable, facetable, retrievable
- Vector fields with dimensions matching the embedding model
- Vector profile: HNSW for speed, exhaustive KNN for exact
- Knowledge store (optional): Projections of enriched data to Blob/Table storage.
- Integrated vectorization = chunking + embedding inside the indexer, and a vectorizer on the index so queries get embedded automatically. Pick it for “least code” RAG ingestion.
Query types
| Type | What it does | Choose when |
|---|---|---|
| Full-text (BM25) | Keyword matching | Exact terms, IDs, codes |
| Vector | Similarity on embeddings | Meaning, paraphrase, multilingual, images |
| Hybrid | Both, merged with Reciprocal Rank Fusion | Default best for RAG |
| Semantic ranker | Re-ranks top results with a language model; captions and answers | Best relevance; needs a semantic configuration |
- Best-practice RAG answer: hybrid + semantic ranker.
- Tune with chunk size/overlap, scoring profiles, filters (security trimming), and query rewriting.
Multimodal ingestion
- Scanned PDFs/images: OCR skill (+ Text Merge) or Document Layout skill, or Content Understanding upstream.
- Images as content: Image verbalization (caption with a multimodal model, embed the text) or multimodal embeddings.
- Audio/video: Transcribe (Speech / Content Understanding), then index the text with timestamps.
Connect retrieval to agents
- Add the Azure AI Search tool to the agent via a project connection.
- Set:
- Index
- Query type (simple, vector, hybrid, semantic, hybrid + semantic)
- top-k
- Filter
- For multi-part questions use agentic retrieval / knowledge bases.
Extract content from documents
- Content Understanding: Multimodal pipeline (OCR + layout + field extraction) driven by an analyzer schema. The AI-103 favorite.
- Outputs markdown (clean for RAG and agents)
- Outputs structured JSON fields with confidence and source grounding
- Document Intelligence: Still valid when a prebuilt model matches exactly.
- Prebuilt models: invoice, receipt, ID, layout, read
- Custom template/neural models
- Low confidence on a field → route to human review.
If the question says X, pick Y
Keywords in the question stem usually point straight at one answer. Read this the night before and the morning of.
| Question stem says | Answer is usually |
|---|---|
| “Data must stay in the EU/US” + pay per token | Data Zone Standard deployment |
| “Predictable latency”, “consistent high throughput” | Provisioned (PTU) deployment |
| “Millions of docs overnight”, “lowest cost”, “not time-sensitive” | Global Batch |
| “429 Too Many Requests” | Retry with backoff + retry-after; raise quota; add deployments behind APIM |
| “Without storing keys”, “most secure auth” | Managed identity + Entra ID (DefaultAzureCredential), disable local auth |
| “No public internet access” | Private endpoints + disable public network access (+ Standard agent setup with BYO resources) |
| “Agent answers from a few PDFs, least effort” | File Search tool |
| “Agent answers from existing enterprise index” | Azure AI Search tool |
| “Latest news”, “current prices” | Grounding with Bing / web search tool |
| “Compute”, “analyze a CSV”, “make a chart” | Code Interpreter |
| “Existing REST API with OpenAPI 3 spec” | OpenAPI tool |
| “Reuse tools across agents via an open protocol” | MCP tool |
| “Human must approve before the action” | require_approval on the tool / approval step in the workflow |
| “Orchestrator delegates to specialist agents” | Connected agents (or a handoff/sequential workflow) |
| “Answer invents facts”, “fabrication” | Groundedness evaluation + groundedness detection; improve retrieval |
| “User tries to jailbreak” | Prompt Shields (user prompt attacks) |
| “Malicious instructions hidden in a document/email/image” | Prompt Shields (indirect/document attacks); treat content as data |
| “Output reproduces song lyrics or licensed code” | Protected material detection |
| “Reliable JSON matching a schema” | Structured outputs (JSON schema) |
| “Deterministic, less creative output” | Lower temperature |
| “Best relevance for RAG” | Hybrid search + semantic ranker |
| “Least code to chunk and embed on ingest” | Integrated vectorization (Text Split + embedding skill + vectorizer) |
| “Scanned PDFs in the index” | OCR skill + Text Merge, or Document Layout skill |
| “Custom logic during indexing” | Custom Web API skill (Azure Function) |
| “Extract fields + markdown from docs/images/video for agents” | Content Understanding analyzer |
| “Cross-document reasoning, validate against reference data” | Content Understanding pro mode |
| “Edit only part of an image” | Image edit with a mask (inpainting) |
| “Redact PII at scale” | Azure AI Language PII detection |
| “Translate Word/PDF, keep formatting” | Translator document translation |
| “Speech misrecognizes product names” | Phrase list first, then custom speech |
| “Change pronunciation, pauses, voice style” | SSML |
| “Low-latency voice agent” | Voice Live API / realtime audio model |
| “Trace latency per tool call and token usage” | OpenTelemetry tracing to Application Insights |
| “Run evals automatically before deploy” | Evaluation step in CI/CD (GitHub Actions / Azure DevOps) |
Common traps
- Function calling does not run your code. The model returns the call; your app executes it and submits the output. The OpenAPI and MCP tools are the ones the service calls for you.
- File Search vs Azure AI Search tool:
- File Search = you upload files, the service builds the vector store
- AI Search tool = you already own and manage an index
- Hybrid is not semantic. Hybrid = keyword + vector merged by RRF. Semantic ranker is a separate re-ranking layer on top.
- Vector field dimensions must match the embedding model’s output size. A mismatch breaks indexing.
- Temperature and top_p: Tune one, not both.
- Fine-tuning does not add fresh knowledge well.
- Knowledge gaps → RAG
- Style/format/behavior → fine-tuning
- Global Standard is not a residency answer. Residency → Data Zone or Regional.
- API keys are never the “most secure” answer when managed identity is offered.
- Content filters are not just output filters. They check prompts and completions. Prompt Shields handle jailbreaks and indirect attacks specifically.
- Old names in answers: “Azure OpenAI Service”, “Azure AI Studio”, “hub-based project”, “Cognitive Services” may appear.
- Map them to Foundry resource/project/Tools
- Prefer the Foundry-native option when both are listed
Self-check questions
| Question | Answer |
|---|---|
| Your agent must summarize incoming emails, and attackers embed instructions in email bodies. What detects it? | Prompt Shields, indirect (document) attack detection |
| You need 99th-percentile latency guarantees for a chat app with steady load. Deployment type? | Provisioned (PTU) |
| Agent must look up order status from your internal service that runs in your app process. Tool? | Function calling |
| RAG answers cite the wrong chunks. First fix? | Hybrid search + semantic ranker, then tune chunk size/overlap |
| Which evaluator measures whether answers are supported by retrieved context? | Groundedness |
| Agent threads and files must be stored in your own Cosmos DB and Storage accounts. Setup? | Standard agent setup (bring your own resources) |
| Extract invoice fields and a markdown version of each invoice for an agent, validating totals against a price list. | Content Understanding, pro mode with reference data |
| Replace the sky in a product photo but keep everything else. | Image edit with a mask (inpainting) |
| Call center STT keeps missing drug names; no training data yet. | Phrase list |
| A deployed agent is slow; you need to see which tool call takes longest. | Tracing (OpenTelemetry → Application Insights), inspect spans |
| Which Search query type merges BM25 and vector results? | Hybrid (Reciprocal Rank Fusion) |
| Least-code way to chunk and embed blobs during indexing? | Integrated vectorization |