Index

Exam overview

AI-103 is Foundry-first. More than half the score (55-65%) is planning/securing Foundry and building gen AI apps and agents, so most of the study time should go there.

  • Pass mark is 700/1000.
  • Questions are scenario-based, Python-flavored, and mostly on GA features (common preview features can appear).
  • The shift from AI-102: stop thinking “which Cognitive Service?” and start thinking “which Foundry capability, model, or tool?”
  • The old services now show up as Foundry Tools (Language, Speech, Translator, Vision, Content Understanding, Document Intelligence) inside a Foundry resource + project.

Exam domains

Domain Weight What it covers
Plan and manage an Azure AI solution 25-30% Model/service choice, deployment types, quotas, security, responsible AI
Generative AI and agentic solutions 30-35% RAG, Foundry Agent Service, tools, multi-agent, evals, tracing
Computer vision 10-15% Image/video generation and editing, multimodal understanding, Content Understanding
Text analysis (incl. speech) 10-15% LLM extraction, Language, Translator, Speech
Information extraction 10-15% Azure AI Search, skillsets, Content Understanding for documents

If you are coming from AWS

Azure / Foundry Closest AWS mental model
Foundry resource + project Bedrock account setup + a workspace boundary
Foundry Agent Service Bedrock Agents / AgentCore Runtime
Foundry Tools catalog, MCP tool, OpenAPI tool Action groups, AgentCore Gateway
Azure AI Search OpenSearch / Kendra / Bedrock Knowledge Bases
Azure AI Content Safety, guardrails Bedrock Guardrails
Content Understanding Bedrock Data Automation
Foundry evaluations + tracing (App Insights) Bedrock evaluations + AgentCore Observability

Exam day tips

  • Read the last sentence of the question first. It tells you what is being asked (cheapest, least effort, most secure).
  • Flag the question and move on if it is taking more than 90 seconds.
  • Case studies can’t be revisited once you leave them.

Plan and manage an Azure AI solution

Most questions here are “pick the right option under a constraint” (cost, latency, residency, security). Know the object model, the deployment types, and the security defaults cold.

Foundry object model

  • Foundry resource (kind AIServices): The top-level Azure resource. Holds model deployments, Foundry Tools, networking, keys/identity.
  • Foundry project: The working boundary inside the resource. Holds agents, connections, evaluations, traces, files.
    • Apps connect with the project endpoint: https://<resource>.services.ai.azure.com/api/projects/<project>
    • Use AIProjectClient + DefaultAzureCredential
  • Connections: How a project reaches outside resources (Azure AI Search, Storage, Bing grounding, MCP servers, APIs).
    • Credentials live in the connection, not in code or prompts.

Choosing a model

Need Pick
General chat, strong reasoning, tool calling Flagship LLM (GPT-4.1 / GPT-5 class)
Hard multistep reasoning, math, planning Reasoning model (o-series / GPT-5 reasoning)
Cheap, fast, edge or offline, simple tasks Small language model (Phi family, mini/nano variants); Foundry Local for on-device
Images + text in one call Multimodal model (GPT-4o / 4.1 class)
Vectors for search Embedding model (text-embedding-3-small/large)
Generate or edit images gpt-image-1 class, FLUX
Generate video Sora class
Low-latency voice conversation Realtime / audio models, or Speech + LLM
Prebuilt task (translate, OCR, PII, STT) A Foundry Tool, not an LLM
  • The smallest model that meets quality wins cost questions.
  • A Foundry Tool beats a prompt when the task is standard and deterministic output matters.
  • Model names change often. Where a model is named, treat it as “the current model of that class”.

Deployment types

Type Use when
Global Standard Default. Pay per token, highest quota, data may process in any Azure region
Data Zone Standard Pay per token but processing must stay in the US or EU data zone
Standard (regional) Processing must stay in one region
Provisioned (Global / Data Zone / Regional), PTUs Predictable latency and throughput for steady high volume; reserved capacity
Global Batch / Data Zone Batch Large async jobs, about 50% cheaper, results within 24 hours
Serverless API vs managed compute Partner/open models: serverless = pay per token, managed compute = you pay for VMs

Quotas, scale and cost

  • Quota is TPM (tokens per minute) and RPM per model, per region, per subscription.
  • Hitting the quota returns HTTP 429.
    • Honor the retry-after header
    • Use exponential backoff
  • Scale beyond one deployment: more regions/deployments behind Azure API Management as an AI gateway.
    • Load balancing
    • Token-limit policy
    • Token metrics
    • Semantic caching
  • Provisioned can spill over to Standard.
  • Cost levers:
    • Smaller model
    • Batch
    • Prompt caching
    • Shorter prompts / max_tokens
    • Semantic cache
    • PTU reservations for steady load

Monitoring

  • Models/agents:
    • Azure Monitor metrics (tokens, latency, 429s)
    • Application Insights tracing via OpenTelemetry
    • Foundry observability dashboards
    • Continuous evaluation on production traffic
  • Search:
    • Indexer execution history and errors
    • Index size / document count
    • Query latency
    • Relevance testing

Security

  • Keyless: Microsoft Entra ID + managed identity, DefaultAzureCredential, disable local (key) auth. Keys only in Key Vault if you must.
  • RBAC: Least privilege.
    • A user/app role to call models and agents (Azure AI User / Cognitive Services OpenAI User)
    • A higher role to create deployments and manage the resource
    • Grant the Foundry managed identity roles on the AI Search and Storage it reads
  • Network: Private endpoints, disable public network access, VNet injection for agents.
  • Agent setup:
    • Basic = Microsoft-managed storage
    • Standard = bring your own Cosmos DB (threads/conversations), Storage (files), AI Search (vectors)
    • Pick Standard for compliance, residency, or private networking
  • CI/CD: Azure Developer CLI (azd) + Bicep/Terraform, GitHub Actions or Azure DevOps. Run evaluations as a pipeline gate before promotion.

Responsible AI

  • Content filters / guardrails: Applied to input and output.
    • Harm categories: hate, sexual, violence, self-harm
    • Severity levels: safe, low, medium, high
  • Add-ons:
    • Prompt Shields = user jailbreaks + indirect/document attacks
    • Groundedness detection
    • Protected material (text and code)
    • Custom blocklists
    • PII detection
    • Task adherence for agents
  • Evaluators:
    • Quality: groundedness, relevance, coherence, fluency, similarity, retrieval
    • Safety: harmful content, indirect attack, protected material, code vulnerability
    • Agent: intent resolution, tool call accuracy, task adherence
  • AI Red Teaming Agent (PyRIT) for adversarial testing.
  • Auditing: Trace logs, provenance metadata (C2PA content credentials on generated images), approval workflows.
  • Agent governance:
    • Tool allow-lists
    • Per-tool auth
    • require_approval="always" on MCP tools
    • Human-in-the-loop for high-impact actions
    • Agent identity in Entra

Generative AI and agentic solutions

The biggest domain. If you only master one thing: which agent tool solves which problem, and how grounding (RAG) is wired into agents.

RAG

  • Ingest → chunk → embed → index (Azure AI Search)
  • At query time, run hybrid search (keyword + vector) + semantic ranker
  • Put the top chunks in the prompt
  • Answer with citations
  • Evaluate groundedness
  • Agentic retrieval (Azure AI Search knowledge bases / Foundry IQ): Newer pattern where an LLM plans and runs subqueries for you. Use it for complex multi-part questions.

Agent building blocks

  • Agent = model + instructions (role, goal, constraints) + tools.
  • Kinds of agents:
    • Prompt agent: declarative, no code hosting
    • Workflow: multi-step orchestration, visual or YAML
    • Hosted agent: your containerized code (e.g., Microsoft Agent Framework or LangGraph)
  • Conversation state:
    • New API uses conversations + Responses API
    • Classic API used thread → message → run
    • Classic agents are deprecated and retire March 31, 2027, but exam questions may still show thread/run code. Know both vocabularies.
  • Memory:
    • Short-term = the conversation
    • Long-term = agent memory store (preview) or your own store (Cosmos DB)
    • Trim or summarize long histories to fit context
  • Tool schemas: Function tools are JSON Schema (name, description, parameters, required). Good descriptions drive correct tool selection.

Agent tools

Scenario Tool
Answer from a few uploaded files, zero infra File Search (managed vector store)
Answer from an existing enterprise index Azure AI Search tool (via project connection)
Current public info from the web Grounding with Bing / Web search
Math, data analysis, charts, file transforms Code Interpreter (sandboxed Python)
Call your own logic in your app process Function calling (model returns a call; your code runs it and submits output)
Call an existing REST API with a spec OpenAPI tool (anonymous, API key via connection, or managed identity)
Reuse a remote tool server MCP tool (server_label, server_url, require_approval, connection for auth)
Low-code business workflow Azure Logic Apps / Azure Functions
M365 docs or Fabric data SharePoint tool / Fabric data agent
Multi-step web research report Deep Research tool
Let a main agent delegate Connected agents / A2A

Multi-agent orchestration

  • Connected agents: An orchestrator agent calls specialist agents as tools. Simple delegation, no custom code.
  • Workflows / Microsoft Agent Framework patterns:
    • Sequential
    • Concurrent (fan-out/fan-in)
    • Group chat
    • Handoff
    • Human-in-the-loop checkpoints
  • Agent Framework is the successor to Semantic Kernel + AutoGen.
  • A2A protocol is for agents across platforms; MCP is for tools.
  • Safeguards for autonomy:
    • Approval steps before irreversible actions
    • Max iterations / turn limits
    • Tool allow-lists
    • Content filters on both input and output
    • Task adherence checks

Code shape to recognize

from azure.identity import DefaultAzureCredential
from azure.ai.projects import AIProjectClient

project = AIProjectClient(endpoint=PROJECT_ENDPOINT,
                          credential=DefaultAzureCredential())
# new API: define agent (model + instructions + tools),
# then call it through the OpenAI-compatible Responses API
# classic API: create_agent -> threads.create -> messages.create -> runs.create_and_process
  • When code uses an API key where Entra would work, the “most secure” answer swaps in DefaultAzureCredential.

Evaluations

Question Evaluator
Fabrication / hallucination Groundedness (is the answer supported by retrieved context?)
Did retrieval work? Retrieval / relevance
Is it well written? Coherence, fluency
Agents Intent resolution, tool call accuracy, task adherence
  • Run with the azure-ai-evaluation SDK or the Foundry portal, on a test dataset.
  • Compare runs, wire into CI/CD, enable continuous evaluation in production.
  • Error analysis = read the traces of failed cases.

Optimize and operationalize

  • Order of fixes: prompt engineering → RAG (knowledge gaps) → fine-tuning (style/format/behavior, or distill to a cheaper model)
  • Fine-tuning types: SFT, DPO (preference), RFT (reasoning models)
  • Parameters:
    • temperature or top_p = change one, not both
    • Low temperature for extraction/factual, higher for creative
    • max_tokens, stop sequences, frequency/presence penalties
    • reasoning_effort on reasoning models
    • Structured outputs (JSON schema) for reliable JSON
  • Prompting: Clear system message, few-shot examples, delimiters, ask for citations, chain-of-thought for multistep tasks.
  • Reflection / self-critique: generate → critique (same or judge model) → revise, with a stop condition. LLM-as-judge is how evaluators work.
  • Observability: OpenTelemetry tracing to Application Insights. Spans per model call and tool call show token usage, latency breakdown, and safety signals.
  • Hybrid orchestration:
    • Route by task (small model for classification, big model for reasoning)
    • Use a rules engine for deterministic compliance checks and the LLM for language

Computer vision

Three buckets: generate/edit media, understand media with multimodal models or Content Understanding, and keep it safe.

Generate and edit

  • Images (gpt-image-1 class, FLUX):
    • Text-to-image
    • Image + reference images
    • Edits with a mask (inpainting): the transparent area of the mask PNG is what gets regenerated
    • Controls: size, quality, number of images, background (transparent), output format
  • Video (Sora class):
    • Text-to-video, image-to-video, and edit/remix an existing video
    • It is an async job: create job → poll status → download
    • Controls: resolution, duration, number of variants
  • Generated media carries C2PA content credentials (provenance metadata) so it can be identified as AI-generated.

Understand images and video

  • Multimodal chat: Pass images as image_url (public URL or base64 data URI) in the message content.
    • detail: low = cheap, fast
    • detail: high = fine detail, more tokens
    • Multiple images in one message for comparison
  • Captions: Concise vs detailed is controlled by the prompt (and max_tokens).
  • Alt text:
    • Short and functional (about one sentence, no “image of”)
    • Decorative images get empty alt
    • Complex images (charts) get an extended description
  • Visual Q&A: Instruct the model to answer only from what is visible and say when it can’t tell.
  • Locate objects/regions: Content Understanding or Azure AI Vision object detection give bounding boxes. Plain LLM chat is weaker at precise coordinates.

Azure Content Understanding

A Foundry Tool where you define an analyzer with a field schema (or use a prebuilt one) for documents, images, audio, or video.

  • Returns structured fields + markdown, with confidence and grounding.
  • Video: segments, shot/scene detection, key frames, transcript, per-segment fields.
Mode Use when
Standard (single-task) One file, one extraction
Pro Multi-file, cross-document reasoning, can use reference data for validation; complex, multi-step extraction

Responsible AI for images

  • Content filters apply to image inputs and outputs (hate, sexual, violence, self-harm).
  • Indirect prompt injection via text in images (a photo containing “ignore previous instructions”):
    • Treat image text as untrusted data
    • Extract it with OCR and scan with Prompt Shields
    • Keep system instructions separate
    • Limit which tools the agent can call
  • Visual policy rules:
    • Watermarks/provenance on generated media
    • Detect prohibited symbols or brand misuse with custom analyzer fields or custom categories
    • Block or route for human review

Text analysis and speech

The recurring question is “LLM prompt or Foundry Tool?”

  • Pick the Tool for standard, repeatable, auditable tasks (PII redaction, sentiment at scale, document translation).
  • Pick the LLM for flexible schemas, nuance, and tone.

Text analysis

  • Structured JSON from an LLM: Use structured outputs (response_format with a JSON schema, strict) instead of “please return JSON”. Low temperature for extraction.
  • Azure AI Language (Foundry Tool):
    • NER
    • PII detection and redaction
    • Key phrases
    • Sentiment + opinion mining
    • Language detection
    • Summarization
    • Custom NER/classification
    • Conversational language understanding (CLU)
    • Question answering
  • Safety and sensitive content: Azure AI Content Safety for harmful text; Language PII for personal data.
  • Domain customization (compliance summaries, domain extraction):
    • System prompt with glossary and rules
    • Few-shot examples
    • Output schema
    • Custom NER or fine-tuning if prompts aren’t enough

Translation

Need Pick
Translate strings in real time, many target languages in one call Azure Translator text translation
Translate whole files, keep formatting Translator document translation (async, Blob Storage source/target, managed identity)
Company terminology Custom Translator (train on parallel data) or glossary
Tone, style, context-aware rewrites LLM-powered translation flow

Speech

  • Speech to text:
    • Real-time (streaming)
    • Fast transcription (synchronous, files)
    • Batch (large volumes, async)
    • Diarization for who-spoke-when
  • Accuracy on domain terms:
    • Phrase list first (no training, quick)
    • Custom speech model (train with text and audio) when accents, noise, or vocabulary need more
  • Text to speech:
    • Neural voices
    • SSML controls pronunciation, pauses, rate, pitch, speaking style
    • Custom neural voice is limited access (approval needed)
  • Speech as an agent modality:
    • Voice Live API (low-latency speech-to-speech, can front a Foundry agent)
    • Realtime audio models
    • The classic STT → agent → TTS pipeline
    • Handle barge-in and turn detection
  • Reasoning over audio: Audio-input models, or Content Understanding audio analyzers (transcript + extracted fields).
  • Speech translation: Speech service translation (speech to text/speech in target languages) or STT → LLM translate → TTS.

Information extraction

This is Azure AI Search plumbing plus Content Understanding for documents. Know the indexer pipeline order and the four query types.

Azure AI Search pipeline

  • Data source: Blob, ADLS, SQL, Cosmos DB, SharePoint, etc.
  • Indexer: Pulls on a schedule, change and deletion detection, field mappings. Check execution history for errors/warnings.
  • Skillset (enrichment):
    • Built-in skills:
      • OCR
      • Image Analysis
      • Document Layout (layout-aware, markdown chunks)
      • Text Split (chunking)
      • Text Merge (put OCR text back into content)
      • Entity recognition, key phrases, language detection
      • Azure OpenAI Embedding
    • Custom skill: Web API skill (usually an Azure Function) with a fixed JSON input/output contract
  • Index:
    • Fields with attributes: key, searchable, filterable, sortable, facetable, retrievable
    • Vector fields with dimensions matching the embedding model
    • Vector profile: HNSW for speed, exhaustive KNN for exact
  • Knowledge store (optional): Projections of enriched data to Blob/Table storage.
  • Integrated vectorization = chunking + embedding inside the indexer, and a vectorizer on the index so queries get embedded automatically. Pick it for “least code” RAG ingestion.

Query types

Type What it does Choose when
Full-text (BM25) Keyword matching Exact terms, IDs, codes
Vector Similarity on embeddings Meaning, paraphrase, multilingual, images
Hybrid Both, merged with Reciprocal Rank Fusion Default best for RAG
Semantic ranker Re-ranks top results with a language model; captions and answers Best relevance; needs a semantic configuration
  • Best-practice RAG answer: hybrid + semantic ranker.
  • Tune with chunk size/overlap, scoring profiles, filters (security trimming), and query rewriting.

Multimodal ingestion

  • Scanned PDFs/images: OCR skill (+ Text Merge) or Document Layout skill, or Content Understanding upstream.
  • Images as content: Image verbalization (caption with a multimodal model, embed the text) or multimodal embeddings.
  • Audio/video: Transcribe (Speech / Content Understanding), then index the text with timestamps.

Connect retrieval to agents

  • Add the Azure AI Search tool to the agent via a project connection.
  • Set:
    • Index
    • Query type (simple, vector, hybrid, semantic, hybrid + semantic)
    • top-k
    • Filter
  • For multi-part questions use agentic retrieval / knowledge bases.

Extract content from documents

  • Content Understanding: Multimodal pipeline (OCR + layout + field extraction) driven by an analyzer schema. The AI-103 favorite.
    • Outputs markdown (clean for RAG and agents)
    • Outputs structured JSON fields with confidence and source grounding
  • Document Intelligence: Still valid when a prebuilt model matches exactly.
    • Prebuilt models: invoice, receipt, ID, layout, read
    • Custom template/neural models
  • Low confidence on a field → route to human review.

If the question says X, pick Y

Keywords in the question stem usually point straight at one answer. Read this the night before and the morning of.

Question stem says Answer is usually
“Data must stay in the EU/US” + pay per token Data Zone Standard deployment
“Predictable latency”, “consistent high throughput” Provisioned (PTU) deployment
“Millions of docs overnight”, “lowest cost”, “not time-sensitive” Global Batch
“429 Too Many Requests” Retry with backoff + retry-after; raise quota; add deployments behind APIM
“Without storing keys”, “most secure auth” Managed identity + Entra ID (DefaultAzureCredential), disable local auth
“No public internet access” Private endpoints + disable public network access (+ Standard agent setup with BYO resources)
“Agent answers from a few PDFs, least effort” File Search tool
“Agent answers from existing enterprise index” Azure AI Search tool
“Latest news”, “current prices” Grounding with Bing / web search tool
“Compute”, “analyze a CSV”, “make a chart” Code Interpreter
“Existing REST API with OpenAPI 3 spec” OpenAPI tool
“Reuse tools across agents via an open protocol” MCP tool
“Human must approve before the action” require_approval on the tool / approval step in the workflow
“Orchestrator delegates to specialist agents” Connected agents (or a handoff/sequential workflow)
“Answer invents facts”, “fabrication” Groundedness evaluation + groundedness detection; improve retrieval
“User tries to jailbreak” Prompt Shields (user prompt attacks)
“Malicious instructions hidden in a document/email/image” Prompt Shields (indirect/document attacks); treat content as data
“Output reproduces song lyrics or licensed code” Protected material detection
“Reliable JSON matching a schema” Structured outputs (JSON schema)
“Deterministic, less creative output” Lower temperature
“Best relevance for RAG” Hybrid search + semantic ranker
“Least code to chunk and embed on ingest” Integrated vectorization (Text Split + embedding skill + vectorizer)
“Scanned PDFs in the index” OCR skill + Text Merge, or Document Layout skill
“Custom logic during indexing” Custom Web API skill (Azure Function)
“Extract fields + markdown from docs/images/video for agents” Content Understanding analyzer
“Cross-document reasoning, validate against reference data” Content Understanding pro mode
“Edit only part of an image” Image edit with a mask (inpainting)
“Redact PII at scale” Azure AI Language PII detection
“Translate Word/PDF, keep formatting” Translator document translation
“Speech misrecognizes product names” Phrase list first, then custom speech
“Change pronunciation, pauses, voice style” SSML
“Low-latency voice agent” Voice Live API / realtime audio model
“Trace latency per tool call and token usage” OpenTelemetry tracing to Application Insights
“Run evals automatically before deploy” Evaluation step in CI/CD (GitHub Actions / Azure DevOps)

Common traps

  • Function calling does not run your code. The model returns the call; your app executes it and submits the output. The OpenAPI and MCP tools are the ones the service calls for you.
  • File Search vs Azure AI Search tool:
    • File Search = you upload files, the service builds the vector store
    • AI Search tool = you already own and manage an index
  • Hybrid is not semantic. Hybrid = keyword + vector merged by RRF. Semantic ranker is a separate re-ranking layer on top.
  • Vector field dimensions must match the embedding model’s output size. A mismatch breaks indexing.
  • Temperature and top_p: Tune one, not both.
  • Fine-tuning does not add fresh knowledge well.
    • Knowledge gaps → RAG
    • Style/format/behavior → fine-tuning
  • Global Standard is not a residency answer. Residency → Data Zone or Regional.
  • API keys are never the “most secure” answer when managed identity is offered.
  • Content filters are not just output filters. They check prompts and completions. Prompt Shields handle jailbreaks and indirect attacks specifically.
  • Old names in answers: “Azure OpenAI Service”, “Azure AI Studio”, “hub-based project”, “Cognitive Services” may appear.
    • Map them to Foundry resource/project/Tools
    • Prefer the Foundry-native option when both are listed

Self-check questions

Question Answer
Your agent must summarize incoming emails, and attackers embed instructions in email bodies. What detects it? Prompt Shields, indirect (document) attack detection
You need 99th-percentile latency guarantees for a chat app with steady load. Deployment type? Provisioned (PTU)
Agent must look up order status from your internal service that runs in your app process. Tool? Function calling
RAG answers cite the wrong chunks. First fix? Hybrid search + semantic ranker, then tune chunk size/overlap
Which evaluator measures whether answers are supported by retrieved context? Groundedness
Agent threads and files must be stored in your own Cosmos DB and Storage accounts. Setup? Standard agent setup (bring your own resources)
Extract invoice fields and a markdown version of each invoice for an agent, validating totals against a price list. Content Understanding, pro mode with reference data
Replace the sky in a product photo but keep everything else. Image edit with a mask (inpainting)
Call center STT keeps missing drug names; no training data yet. Phrase list
A deployed agent is slow; you need to see which tool call takes longest. Tracing (OpenTelemetry → Application Insights), inspect spans
Which Search query type merges BM25 and vector results? Hybrid (Reciprocal Rank Fusion)
Least-code way to chunk and embed blobs during indexing? Integrated vectorization

Resources