Claude Certified Developer Foundations: Takeaways
Notes from the Claude Certified Developer Foundations prep path (five modules). Writing code that uses Claude is different from using Claude to write code. Each section below covers the teaching for that module, then practice questions from its quizzes and checkpoints with the answers the course marks correct. Confirm current model names, beta headers, and API defaults on Anthropic docs at build time; products move.
Model fundamentals (Module 1)
About an hour of material. Goal: share a vocabulary for tokens, context, sampling, model tiers, prompting modes, and how you call the API before the rest of the path builds on it.
Tokens and the context window
Claude reads tokens, not characters or words. Characters per token depend on the tokenizer. Everything counts toward price and the context budget: prompt, history, tool definitions, tool results, and the reply.
The context window is a fixed token budget for one request. It holds the system prompt, conversation, injected documents, tool results, and model output. If the input is already too large, the request is rejected before generation. If generation hits the ceiling mid-reply, you get truncated text with stop reason model_context_window_exceeded. The API does not silently drop oldest turns; your app must trim or summarize history. Short tests rarely fill the window. Production sessions do sooner.
Sampling and non-determinism
At each step the model samples the next token from a probability distribution. Temperature shapes that distribution (lower is more repeatable, higher is more varied). The same prompt can return different wording. Newest Claude models may reject non-default sampling params (temperature, top_p, top_k) with a 400; steer with prompting instead when that applies. Temperature 0 is more repeatable but not guaranteed identical.
Identical inputs do not guarantee identical outputs. Do not assert exact response text in tests. Assert properties that must hold, or use an eval with a model-graded judge when you need to judge meaning.
Model family and reasoning mode
Capability tiers named on screen: Fable, Opus, Sonnet, Haiku. Tradeoffs of cost, latency, and capability. Sonnet is the balanced default. Haiku favors speed and cost when the task fits. Opus for harder work above Sonnet. Fable for the most demanding reasoning, coding, and agent work. Start with Sonnet. Move up only when an eval shows you miss the quality bar. Move down to Haiku only when an eval shows the quality drop is acceptable.
Reasoning mode is separate from model choice. On current models it is adaptive thinking: the model decides when and how much to think; you tune depth with an effort setting. Older budget_tokens is deprecated and can return 400 on newest models. Thinking content is omitted from responses by default on newest models unless you request summarized display. Useful on hard multi-step work; wasted on lookups and classification.
The two levers combine: capable model with reasoning off is fast and direct; smaller model with reasoning on spends more tokens to think; hardest tasks pair a capable model with higher effort.
Prompting modes
Zero-shot: instruction only, no examples. One-shot: one input/output example. Multi-shot (few-shot): several examples. Examples are not training data; they sit in the prompt to show the answer shape. Each example costs tokens on every call. Use zero-shot when the shape is obvious. Use one-shot or multi-shot when structure, casing, or edge cases keep missing. Prefer the smallest prompt that works. Pair model choice with example count: simplest model and fewest examples that meet the eval.
How you reach Claude
Claude is reached over HTTP REST. Official SDKs wrap the same API and handle auth, request building, retries, and parsing. Synchronous: wait for the full reply. Streaming: pieces arrive as server-sent events; your code reassembles them. High volume: non-blocking clients (for example AsyncAnthropic in Python). Message Batches API: submit many requests, poll for completion (up to about 24 hours), lower cost per token for offline bulk work.
Practice questions
Question: You send the same prompt twice. Do you get identical wording?
Answer: Not necessarily. The model samples each next token, so wording can vary. Prefer property checks or a model-graded eval over exact-string asserts.
Question: How do model choice and reasoning mode relate?
Answer: They are separate levers. Picking a model and turning on adaptive thinking (effort) are different decisions.
Question: Zero-shot already meets the quality bar. What if you add three few-shot examples?
Answer: You mainly add token cost on every call, with little or no gain.
Question: Thousands of offline jobs, lowest cost, no user waiting. What API pattern?
Answer: Message Batches: submit, then poll. Lower cost per token than sync loops.
Question: Input already exceeds the context window vs generation hits the ceiling mid-reply?
Answer: Too-large input is rejected before generation. Mid-generation stop returns what was produced with model_context_window_exceeded. The app must trim or summarize; the API does not silently drop oldest turns.
Production prompting, tools, and agents (Module 2)
Roughly three-plus hours. Writing code that uses Claude is different from using Claude to write code. Topics map to failure modes that are easy to miss in development and costly later: production prompts, extended thinking, tool schemas, streaming, context engineering, agents, memory, multimodal, and batch.
Prompting craft
A prompt that works once in chat often breaks on untested production inputs. Fix by finding the missing structural piece, not by only adding words.
Four techniques:
- System prompts: persistent behavioral contract for the session (role, output format, rules that must not drift).
- XML tags: separate mixed inputs and instructions with clear tags such as
<my_code>and<docs>. Descriptive names are fine. - Few-shot examples: show the format with correct input/output pairs in consistent tags; can reuse strong eval outputs.
- Output constraints: exact field names, types, length limits, preamble rules, and what to do when data is missing. Use structured outputs (JSON schema / strict tool use) when the format must be machine-readable.
Watch out: longer is not better. Quiet failures often mean constraints were not precise enough. A multi-pass story on screen shows adding more words without hard constraints: parsers break on sentences, case drift, multi-labels, and long unfocused replies. The fix that worked was a tight output rule plus few-shot examples. Diagnosing the wrong problem and making the prompt verbose are separate failure modes.
Extended thinking
Shapes how much work Claude does before answering. Thinking comes back as its own block ahead of the answer. Adaptive thinking: enable where needed; tune depth with effort (not deprecated budget_tokens). Thinking tokens cost like output tokens. Use for multi-step constrained reasoning and agent planning; leave off for classification and lookups.
Carry-back rule: return thinking blocks to the API unchanged (signature). Editing, summarizing, or dropping them rejects the next request. Redacted thinking blocks follow the same rule. If context pressure from thinking is the worry, fix it with context engineering, not by stripping blocks.
Tool-use and schema design
Claude does not run tools. It reads your schemas, picks a tool, and returns a tool_use block (name, unique id, arguments). Your app runs the tool and sends a tool_result with the same id in the immediately following user turn. Missing, mismatched, or late results fail API validation. Preserve the full assistant content array (including any text beside tool_use) when appending history.
Schema parts: name, description, input_schema. The description drives selection. Write when to use and when not to use. Overlapping "find information" descriptions cause wrong picks; add exclusion sentences to both tools, or merge them with a type parameter. MCP can supply maintained tools; still tune descriptions and allowlist tools you actually want in the pool.
Streaming
Events build the message: message_start, content_block_start, deltas, content_block_stop, message_delta, message_stop. Do not parse tool JSON until the block stops. Append to history only after message_stop. On interrupt, discard the partial turn and retry. A stream ending is not a message completing. Check stop_reason from message_delta before continuing a loop (tool_use means assembled calls are ready to run). Prefer non-streamed calls for short backend jobs where nobody is waiting.
Watch out: handlers that append whenever the read loop ends will commit half-built tool_use JSON after a network blip. The failure shows up on the next request as a validation error, so teams often debug the schema instead of the stream gate.
Context engineering
Model choice sets cost and latency floors. Then the context window holds everything for the turn. Tool results stay and grow. Production tool outputs are often several times fixture size, so a session that finished in testing can hit a team budget at turn eight. Symptoms look like bad tool selection because instructions were crowded out.
Strategies: pruning (rewind and drop unproductive stretch), compaction (summarize to keep working on the same feature), clearing (new session when the next task is unrelated), subagent handoffs (isolated window, return a summary). Measure production-sized tool outputs before ship. When tool selection degrades after a fixed turn count, check the window before rewriting schemas.
Agent construction
An agent is a multi-step tool-use loop with managed context and a defined goal. Choose a workflow when you can enumerate exact steps in code. Choose an agent when you can specify the goal and tools but not the path. Wrong pattern choice only surfaces in production.
Wiring paths (who runs the loop): raw Messages API (you own everything), Agent SDK (loop in your process; set settingSources explicitly for filesystem skills/CLAUDE.md), Claude Managed Agents (hosted loop and sandbox; beta). Put a human-in-the-loop gate before irreversible writes. "Validation passed on this file" is not enough if downstream systems depended on the old value. Ask at design time: what is the worst outcome if this tool runs unchecked?
Agent memory
Match scope to session shape:
- In-context: short single session; state dies when the session ends.
- External storage: continuity across days or users; adds retrieval latency and read/write logic.
- Summarized memory: long conversations under a budget; loses anything the summarizer dropped.
- Stateless: one-shot jobs that finish and close.
Skills carry repeatable instructions on demand by description match. Subagents do not auto-inherit skills unless listed. Designing memory under production pressure is expensive; decide at design time.
Multimodal and batch
Images spend visual tokens (roughly ceil(width/28) × ceil(height/28) patches). Measure production image cost before building ingestion. Send once-used images inline (base64). Reuse via Files API file_id. Use stable public URLs only when reachable at request time. PDFs use document blocks with the same source patterns.
Message Batches: high-volume offline submit-and-poll, lower per-token cost, non-deterministic latency (can take hours). Chunking a list and looping the sync API is not batching and still hits rate limits. Match results with custom_id because return order is arbitrary. Do not use batches for user-facing waits.
Practice questions
Question: System prompt is only "Extract the key information" for a JSON ticket with category, urgency, and summary. What is missing?
Answer: Hard output constraints: exact field names, allowed values/enums, and rules for missing data. Adding more vague wording is not the fix.
Question: Match each task to extended thinking:
- Classify 50,000 tickets overnight into three labels.
- Plan a multi-step refactor where each step depends on the last.
- Strip thinking blocks from history to save context before the next tool call.
Answer: (1) Leave it off. (2) Enable it and budget for planning. (3) Never do this. Thinking blocks must return unchanged or the next request fails.
Question: Assistant issued toolu_01, but the next user turn sent tool_result with toolu_02. What broke?
Answer: The tool_use_id must match. Every tool_use needs a matching tool_result in the immediately following user turn.
Question: Stream handler appends the assistant turn when the read loop ends. What is wrong?
Answer: Append only after message_stop. On interrupt, discard the partial turn and retry. A stream ending is not a complete message.
Question: Tool selection gets worse after turn 4 while fetching large policy docs. Schema looks fine. Likely fix?
Answer: Prune or compact tool results. Accumulated outputs crowd out instructions; it often looks like a schema bug.
Question: Match memory scope:
- Support agent across daily check-ins for two weeks.
- One-shot document formatter that exits.
- Multi-hour coding session that will not continue later.
Answer: (1) External storage. (2) Stateless. (3) In-context.
Question: Match encoding / API:
- Same product diagram in every request.
- One-off UI bug screenshot.
- Classify 5,000 feedback responses offline.
Answer: (1) Files API (file_id). (2) Inline base64. (3) Message Batches (submit and poll).
Claude Code, MCP, and integration (Module 3)
Roughly two-plus hours. Claude Code runs the same agent loop in your terminal with a permission layer, durable project config, shareable packaging, and MCP connections. The recurring problem: something that worked on your machine must still be safe and usable when a teammate or auditor runs it elsewhere.
Permission modes and human gates
Claude Code explores (reads and traces), then plans (structured intended edits), then codes after you approve. Plan mode can hold the agent in explore so it proposes without writing.
Modes trade speed for oversight:
- default: reads only auto-approved; prompts before nearly every edit or command. Safe baseline; slow on trusted work.
- acceptEdits: auto-approves reads, file edits, and common filesystem commands (
mkdir,touch,rm,rmdir,mv,cp,sed) inside the working directory. Protected paths still prompt. Other shell and out-of-tree writes still gate. Trusted local refactors; not for free-running scripts. - plan: reads only until you approve a plan. Sensitive or unfamiliar codebases.
- auto: classifier reviews actions and blocks escalations / hostile patterns; research preview, not a full substitute for review of sensitive ops.
- dontAsk: only allow-listed tools plus read-only; everything else denied. Locked-down CI/scripts.
- bypassPermissions: all tools, no normal checks. Only catastrophic deletes like
rm -rf /still prompt. Only for disposable isolated environments. Never on a live workstation.
Settings levels: user (~/.claude/settings.json), project (.claude/settings.json, committed), local (.claude/settings.local.json, git-ignored), enterprise (managed-settings.json, not overridable). Deny always wins over allow, including under bypass. Enterprise deny is the durable governance control.
Human gates answer: what is the worst outcome if this runs unchecked? Let low-stakes reversible edits through. Gate hard-to-undo or sensitive actions before they execute. Never let the agent be the only gate on team-marked sensitive code.
Watch out: bypassPermissions to skip "routine" prompts can delete out-of-scope files with no confirmation (course story: pattern matched /src/ and /deploy/config/prod/). Prefer classifier-gated modes and deny rules on sensitive directories before loosening prompts. Note: acceptEdits can silently allow in-tree rm; default would still prompt.
Durable project context
CLAUDE.md at the project root loads every session. /init can generate a starter; keep it to constraints that change behavior. Size dilutes rules and burns context. Move the rest into Skills.
Rules in .claude/rules/ scope with a paths glob in YAML frontmatter. Without paths, they load like CLAUDE.md. Folder layout under rules is organizational only.
Hooks run your scripts at lifecycle points independent of the model. PreToolUse can exit code 2 and write stderr to block a call. PostToolUse cannot block (formatters, tests, audit). Also UserPromptSubmit, Stop, Notification, SessionStart, SessionEnd. A PreToolUse path block is a guardrail, not a CLAUDE.md wish.
Subagents: isolated context, return a summary. Explore/Plan skip CLAUDE.md and git status; general-purpose loads both. Custom subagents do not auto-inherit skills; list them in front matter.
Watch out: CLAUDE.md past hundreds of lines (course example 800+) can contain a path restriction that still gets ignored. Dilution. Path-specific rules go to rules files; history to on-demand docs; the critical rule should be backed by a hook.
Packaging workflows
Skills (SKILL.md under .claude/skills) are portable procedures. Same file, different load paths:
- Claude Code: filesystem discovery by description or name.
- Messages API: sent with the request in Anthropic's code-execution container; needs code-execution and skills beta headers; no local FS assumptions.
- Agent SDK: set
settingSources/setting_sourcesexplicitly or skills may never load. - Managed Agents: skill on the agent resource; Anthropic sandbox; managed-agents beta; sessions server-side (module notes ZDR/HIPAA limits).
Portability: clear descriptions; no baked local tools/paths in the body; subagents must list skills. Custom commands / skills with disable-model-invocation: true for explicit-only. Plugin commands are namespaced (/payments:run-tests). Plugins bundle skills, hooks, subagents, MCP for marketplace install; enterprise managed settings can deploy org-wide.
Watch out: install success only copies files. Absolute author paths and undocumented env vars fail for teammates. Use $CLAUDE_PROJECT_DIR and ${CLAUDE_PLUGIN_ROOT}. Document and validate env vars. Test on a clean machine.
MCP servers
MCP separates tool/resource/prompt definitions into a server any MCP client can reuse. Resources are read-only data by address. Prompts are vetted templates by name.
Transport: stdio (local subprocess), HTTP (remote/shared; preferred), SSE (legacy). Claude Code defers tool definitions by default and loads when needed; every connected server still adds to the pool. Prompt caching stores a stable prefix (cache_control ephemeral, up to four breakpoints, exact match, TTL defaults and opt-in 1h, minimum token threshold). RAG / agentic search keep large libraries outside the window and fetch only the needed slice; retrieval quality depends on organized sources.
Scope: local, user, project (.mcp.json committed), enterprise managed. Project-scoped stdio still spawns locally per machine. Permission rules can name mcp__server__tool. Deny on a tool overrides allow on the server. API connector mcp_toolset enabled flags control visibility vs permission-to-run.
Auth examples: GitHub remote MCP with PAT via env header (never inline). Linear-style OAuth for user identity. Secrets only as env/secret-store references.
Watch out: inline keys in committed .mcp.json enter history; later env-var commits do not erase them. Rotate. Pair CLAUDE.md guidance with PreToolUse that blocks inline-looking .mcp.json edits.
Enterprise integration
Production asks who the model acts as, what data it can reach and where processing happens, whether admins can lock config, and whether access is auditable. Auth by type: OAuth for remote user identity; API key via env for service identity; stdio + filesystem permissions and deny rules for local. Separate credentials from config; rotate on schedule and after exposure; scope narrowly; inventory consumers.
Regulated customers care about residency, audit (PostToolUse logs), and locked enterprise config. Code modernization stress-tests explore → plan → code, plan mode before writes, hooks on sensitive paths, and CLAUDE.md for target conventions. Define blast radius, audit coverage, and phase approvals before starting.
Watch out: OAuth redirect URIs are per host/environment. Staging success does not authorize production. Many enterprises require separate OAuth apps per environment. Put registration on the cutover checklist.
Calibrated trust for AI-generated reviews: trust findings proven from the diff; treat runtime claims as hypotheses; put the human gate where a finding becomes hard to reverse.
Practice questions
Question: Trusted local refactor: auto-approve edits, never run destructive shell, never read .env.production. Which two settings pieces?
Options include: default mode; bypassPermissions; allow Bash(npm run:*) with deny on Bash(rm:*) and Bash(git push:*); deny Read(.env.production); allow Bash(*) and Edit(*).
Answer: Allow/deny shell piece (npm run allowed; rm and git push denied) plus deny Read(.env.production). Not bypassPermissions, not blanket Bash(*).
Question: Edits are auto-approved. The agent wants to change a deployment config several production services read. Where does the human gate go?
Answer: A human reviews and approves before the write runs. Wrong values are hard to undo and reach systems outside the file.
Question: Block reads of .env.production with a hook. Lifecycle event and command behavior?
Answer: PreToolUse. Script inspects the tool call and exits with code 2 (reason on stderr) when the path is .env.production. Exit 0 / log-only does not enforce.
Question: Plugin skill fails for teammates after a clean install. Common defect and fix?
Answer: Absolute path baked to the author's machine (for example under /Users/...). Use $CLAUDE_PROJECT_DIR or ${CLAUDE_PLUGIN_ROOT} so paths resolve on any clone.
Question: Agent SDK should load project skills from disk. What config mistake is common?
Answer: Relying on a default for filesystem sources. Set settingSources (or setting_sources) explicitly.
Question: Key was committed inline in .mcp.json, then "fixed" by editing the file. Still wrong?
Answer: Rotate the key and move it to an environment variable. Later commits do not remove secrets from history.
Evals, ops, and security (Module 4)
Roughly three-plus hours. Prove an agent that worked in development still holds under production traffic: evals that define done, tests and tracing that catch regressions, failure handling under rate limits, model and orchestration under cost/latency budgets, and security that survives a regulated review.
Evals and judges
An eval suite is how you know the system is correct when outputs are non-deterministic. Dataset cases need expected behaviors that name the specific required content, not a restatement of the input. Examples from teaching: a meeting transcript summary must list action items with owners; a bug-report extraction must include the bug and repro steps and omit unrelated asides.
Score bands need clear anchors so graders stay consistent. A useful pattern: low means required content is missing; mid means partial coverage; high means complete and faithful to the expected behavior.
When you use LLM-as-judge, calibrate it against human-labeled cases before you trust it at scale. Holdout cases catch regressions after you change prompts, models, or orchestration.
Testing and tracing
Unit, functional, and integration answer different questions. Pieces can all pass while the end-to-end path fails. Those failures often live in the handoff. Course example: retrieve() returns chunk dicts, but build_prompt() never reads content, so the model never sees retrieved text. Fix the seam and add an integration test on that boundary. Do not "fix the parser" or "reword the prompt" when those parts already pass alone. Tracing makes the handoff visible: which tool ran, what shape came back, and what the next step consumed.
Failure handling
Retries are production design. Defect pattern: sleep(0) (or immediate retry) on every exception, including terminal 400-class errors. Instant retries against a rate limit deepen the limit.
Correct pattern: exponential backoff with a cap, honor retry-after when present, retry only retriable errors, fail fast on terminal statuses (for example 400 / 401 / 403 / 404). Log clearly so operators can tell a limit storm from a bad request.
Model selection and routing
Pick from the real constraint, then confirm with the eval bar.
- High-volume classification where Haiku still meets the eval: Haiku (cost at volume).
- Hard multi-step where a wrong early step is expensive and Sonnet misses the bar: Opus.
- Mixed traffic: route with a cheaper default and an Opus override on complex requests. One model either overpays on the easy bulk or underperforms on the hard slice.
Model choice and reasoning mode remain separate levers; here the focus is constraint matching under measured quality.
Cost and orchestration
Match task shape to pattern instead of running every job as a long interactive agent.
- Single-fact lookup on a stable corpus: fetch once, smallest model that meets the eval.
- Broad research with independent parts: fan-out / parallel, then combine.
- User-facing replies that should feel instant: stream.
- Cost-sensitive non-urgent bulk: Message Batches (submit and poll) plus prompt caching when the same context recurs.
Streaming does not lower token price. Batches trade latency for lower per-token cost. Caching only helps when the cached prefix is truly stable.
Security
Layered controls, not a hopeful system prompt:
- PreToolUse hooks that block writes outside an allowed root before the tool runs.
- Explicit deny rules for sensitive paths (for example
/etc, secrets dirs, credential stores). - Credentials from environment variables or a secret store.
- Audit logs of privileged or blocked actions.
Never take write destinations from untrusted fetched content (for example page.suggested_path) without a fixed allowed output root. Course-style fix: write only to a known path such as /workspace/output/summary.txt, enforce with PreToolUse, keep an eval holdout, and use backoff with fail-fast on terminal errors for the API client.
Cumulative failure pattern: missing eval coverage (only a few manual demos), a retry loop that sleeps zero and retries all statuses, and a writer that trusts remote path suggestions with no PreToolUse gate. Fix all three; any one left open fails a reliability or security review.
Practice questions
Question: What makes a good eval expected output?
Answer: Name the specific required content (for example action items with owners). Do not only restate the input.
Question: Unit tests pass. End-to-end fails. retrieve() returns chunk dicts; build_prompt() never reads content. Where is the fix?
Answer: The integration handoff. Align the seam and add an integration test on that boundary.
Question: Retry loop uses sleep(0) and retries every exception, including 400s. What should it do?
Answer: Backoff, honor retry-after, retry only retriable errors, fail fast on terminal 400-class statuses. Instant retries deepen a rate limit.
Question: Match model choice:
- High-volume classification; Haiku meets the eval bar.
- Hard multi-step; wrong early step is expensive; Sonnet misses the bar.
- Mixed traffic: mostly simple, some complex.
Answer: (1) Haiku. (2) Opus. (3) Route: cheaper default with Opus override on complex requests.
Question: Match task to orchestration lever:
- Single-fact lookup on a stable corpus.
- Broad research that splits into independent parts.
- User-facing reply that should feel instant.
- Cost-sensitive, non-urgent batch job.
Answer: (1) Fetch-once with the smallest model that works. (2) Parallel fan-out. (3) Streaming. (4) Message Batches (and caching when context repeats).
Question: Agent write path must survive a security review. Name four controls.
Answer: PreToolUse hook that blocks writes outside an allowed directory; deny rules for sensitive paths; credentials from environment variables; audit of blocked attempts. Never take the write path from untrusted fetched content.
Accelerators and packaging (Module 5)
Roughly two-plus hours. Make a working build survive reuse, contribution, and deployment beyond the engagement that created it.
Packaging for reuse
Hardcoded engagement values break the next team. Classic defect: a review-agent template hardcodes repo_path to one customer so the next engagement edits the loop instead of configuring it. Parameterize engagement-specific values. If reuse requires editing source, it is not packaged yet.
Contributing back
- A focused tool (for example a single API wrapper) belongs in that tool's own repository, with a behavior test.
- A whole customer application does not fit Cookbook-style review as-is. Extract one focused pattern first.
- A fix to an existing Cookbook example goes as a PR in that example's repository. For engagement-sourced code, clear rights and licensing before technical review.
Contribution quality is both technical (tests, focused scope) and legal (permission to share).
Requirements and lifecycle
Separate what must be true from how you build and run it.
Functional requirements are checkable business or process constraints (example: a summary a human must approve before storage). Vague goals like "fast and accurate" are not checkable.
Infrastructure requirements are checkable platform or residency constraints (example: transcript data processed in the EU). Do not confuse residency with prompt-template design, or HITL rules with infrastructure.
Lifecycle placement from teaching:
- Data residency rule → requirements
- Choose Bedrock because the customer's compliance posture lives there → design
- Pin the full model ID, retain the prior pin for rollback, gate promotion on an eval → deploy
- Instrument token cost and latency in production → operate
Testing sits alongside these phases as the proof that keeps promotion honest.
Deployment and versioning
Prefer pinned full model IDs over moving aliases such as a bare opus string. Retain the previous pinned version so you can roll back. Gate promotion on eval results, not on a successful demo alone.
When assembling a constrained enterprise stack, match identity and platform together. Course pattern: Amazon Bedrock, AWS identity reference (not an Anthropic API key in that posture), a pinned full model ID, and the prior pin kept for rollback.
Comparing platforms
Do not choose a platform on laptop familiarity or latency measured from the wrong region. If rejection cites data residency (for example data processed outside the EU), remeasure from the customer region and pick the platform that satisfies that residency requirement. Optimizing latency or adding cache alone does not fix a residency mismatch.
Trust boundaries
When multiple Claude deployments or fetch-then-call pipelines coordinate, mark the seams explicitly. Treat untrusted fetched content as data, not as instructions (treat_as_data). If page or tool output is passed straight into the next call as trusted system text, injection and confused-deputy failures follow. Apply least-privilege read-only across the privileged seam (least_privilege_read_only), not "run as instructions" or full access.
Cumulative packaging defects to fix together: hardcoded repo_path, a moving model alias with no retained prior version, and fetched content treated as trusted instructions. Correct assembly parameterizes the path, pins an approved model ID with rollback retained, wraps fetched content as data under least privilege, and gates promotion on eval.
Practice questions
Question: Review-agent template hardcodes repo_path to one customer. What is the packaging fix?
Answer: Parameterize engagement-specific values so the next team configures the asset instead of editing the loop.
Question: Where does each contribution go?
- Focused API wrapper tool.
- Whole customer app shared as-is.
- One-line fix to an existing Cookbook example from engagement code.
Answer: (1) Its own repo with a behavior test. (2) Extract one focused pattern first. (3) That example's repo, after a rights/licensing check.
Question: Platform was picked on familiarity. Laptop latency looked fine. Rejected: data processed outside the EU. Best fix?
Answer: Remeasure from the customer region and pick the platform that meets EU-only residency. Optimizing laptop latency or adding cache does not fix residency.
Question: Place each decision in the lifecycle:
- Data must be processed in a specific region.
- Choose Bedrock because compliance posture lives there.
- Pin the full model ID and keep the prior version; gate promotion on an eval.
- Instrument token cost and latency in production.
Answer: (1) Requirements. (2) Design. (3) Deploy. (4) Operate.
Question: Fetched page content is passed straight into the next Claude call as if it were trusted instructions. What belongs in the blanks?
Answer: Treat fetched content as data (not instructions), and apply least-privilege read-only across the privileged seam.
Question: Production model string is a moving alias like opus, with no prior version kept. What should deployment look like?
Answer: Pin a full model ID and retain the previous pin for rollback. Gate promotion on an eval result.
Bottom line
Budget in tokens. Test with evals, not exact strings. Keep tool_use and tool_result ids aligned. Do not commit partial streams. Enforce permissions with deny rules and PreToolUse hooks. Match MCP transport and scope to who should load the server, and never commit secrets. Package so the next person can run the work without editing paths from your machine.
Comments