Say Web Solutions

Claude Certified Architect Foundations: Takeaways

Claude AI LLM AI Prompt Engineering

Notes from the Claude Certified Architect Foundations prep path. The path stacks free Skilljar courses so you can practice the same ideas the exam expects: how you call Claude, how you measure prompts, how tools and RAG work, and how MCP fits in. Confirm current model IDs and platform defaults on Anthropic or cloud docs at build time; those names move.

Prep courses covered: AI Fluency, Building with the Claude API, Claude on Google Cloud (Vertex), Claude Code in Action, Claude 101, Claude with Amazon Bedrock, and Introduction to Model Context Protocol. Quiz answers below match what those courses marked correct.

How you talk to Claude

Claude reads tokens, not whole words. Tokenization breaks text into those units, then the model predicts the next one. Temperature changes how adventurous that pick is. Low temperature is steadier for extraction and coding. Higher temperature is more varied for brainstorming.

The API is stateless. It does not remember prior turns unless you send the history again. Streaming shows text as it is generated so users are not staring at a blank screen. System prompts give Claude a role or behavior for the whole conversation. Prefilling an assistant message starts the reply for Claude so you can steer format (for example opening with a JSON fence and a matching stop sequence).

On Bedrock you use the AWS SDK and Bedrock APIs, not the Anthropic API key path. Same Claude models, different client and docs. Inference profiles help when a model is not hosted in every region.

Practice questions

Question: What are tokens in the context of language models?

Answer: The smallest units a language model can understand.

Question: What happens when you increase the temperature setting?

Answer: Claude becomes more creative and varied in responses.

Question: What is the main benefit of using streaming responses?

Answer: Users see text appear immediately instead of waiting.

Question: Why don't multi-turn conversations work automatically with Claude?

Answer: The API doesn't store any previous messages.

Question: What does assistant message prefilling do?

Answer: Provides the beginning of Claude's response to guide its direction.

Prompt evaluation and prompt engineering

Evaluation measures how well a prompt works with objective scores. Engineering is the work of rewriting the prompt to raise those scores. Typical loop: write a first version, run it through test cases and a grader, then improve.

Being clear and direct puts an action verb and the real task in the first line. Being specific adds length, structure, or steps. XML tags mark which text is data versus instructions. Examples (one-shot or multi-shot) show the shape you want, including awkward cases like sarcasm labels.

Graders can be code (keywords, length, syntax) or model-based (score plus strengths, weaknesses, and reasoning). Do not ship after two happy-path runs. Unexpected user inputs will break a prompt that was never stress-tested.

Practice questions

Question: What is the main purpose of prompt evaluation?

Answer: To objectively measure and improve prompt effectiveness.

Question: After writing your first version of a workout-plan prompt, what should you do?

Answer: Test it, see how well it works, then improve it.

Question: You're asking an AI to analyze a long customer review mixed in with your instructions. What helps the AI understand which part is the review?

Answer: Put the review between XML tags like <review></review>.

Question: You want an AI to write a book summary. Which opening instruction works best?

Answer: Write a three-paragraph summary of this book.

Question: You tested a prompt twice then deployed it. What is the main risk?

Answer: Users will provide unexpected inputs that break the prompt.

Tool use

Tool use lets Claude ask your server for live data or actions beyond training data. Flow: Claude requests a tool, your server runs code, you return results, Claude finishes the answer. Tool choice can force a named tool for extraction. Schemas need clear descriptions of the tool and each parameter. Batch tool use helps when several tools should run in one response instead of a slow chain.

Practice questions

Question: When using tools, what happens right after Claude asks for specific external data?

Answer: Your server runs code to fetch the requested information.

Question: You want to force Claude to use a specific tool for data extraction. Which toolChoice setting should you use?

Answer: {"toolChoice": {"tool": {"name": "tool-name"}}}.

Question: You're writing a tool function for Claude. What's the most important thing to include when creating the JSON schema?

Answer: Detailed descriptions of what the tool does and its parameters.

Question: Claude wants to use multiple tools in a single response. What feature allows this?

Answer: Batch tool use.

Question: You're building a chatbot that needs current weather data. What feature helps Claude access it?

Answer: Tool use to call weather APIs.

RAG

Retrieval Augmented Generation breaks large documents into chunks, embeds them, stores vectors, then pulls only the relevant pieces into the prompt for each question. Vector databases store and search those embeddings. BM25 helps with exact IDs and keywords that pure semantic search misses. Contextual retrieval adds document context to each chunk before indexing so chunks do not float without meaning.

Practice questions

Question: You have an 800-page financial report and want to ask an AI specific questions about it. What does RAG help you do?

Answer: Send only the relevant sections to the AI for each question.

Question: What is a vector database in the context of RAG systems?

Answer: A specialized database optimized for storing, comparing, and searching through numerical embeddings.

Question: You're searching for a specific incident ID and semantic search is weak. What works better?

Answer: BM25 lexical search for exact keyword matching.

Question: You send text to an embedding model. What do you get back?

Answer: A list of about 1024 numbers representing the meaning.

Question: What is contextual retrieval?

Answer: A technique that adds context to document chunks before storing them to improve search accuracy.

Features: caching, thinking, images

Prompt caching reuses computational work when the same long prefix appears again. Content before the cache point must be identical, and you need enough tokens (course minimum: 1024). Extended thinking helps after prompt work still leaves accuracy short; it costs more tokens and latency. For images, prompt engineering (steps, methods, examples) matters more than dumping more pictures.

Practice questions

Question: You send Claude the same long document twice in a row. What does prompt caching help with?

Answer: It saves the computational work from processing the text in the document.

Question: You want to cache a short message that's 500 tokens long. What will happen?

Answer: It won't be cached because it's too short.

Question: You've optimized your prompt but Claude still isn't accurate enough on a complex task. What should you consider next?

Answer: Use extended thinking to improve accuracy.

Question: What is an effective technique for increasing Claude's effectiveness with images?

Answer: Using prompt engineering techniques.

MCP

Model Context Protocol is a client/server way to give Claude tools, resources, and prompts without hand-writing every integration schema. Tools are model-controlled. Resources are app-controlled (good for @document style UI data). Prompts are user-controlled (buttons and slash commands). The Python SDK defines tools with @mcp.tool. Test servers with the MCP Inspector (mcp dev ...). Same-machine setups often talk over standard input/output. Templated resources use URI parameters such as docs://documents/{doc_id}.

Practice questions

Question: Without MCP, what's the main problem building a chat app over GitHub data?

Answer: You'd have to write and maintain all the GitHub tool functions yourself.

Question: What's the easiest way to define a new MCP tool with the Python SDK?

Answer: Use the @mcp.tool decorator on a function.

Question: Claude automatically decides to use a calculator tool. Who is controlling that tool usage?

Answer: Claude (the AI model) itself.

Question: Users click a "Format Document" button. Which MCP primitive is that?

Answer: Prompts, because the user directly triggered it.

Question: You want a resource that fetches different documents by ID in the URI. What type should you use?

Answer: A templated resource with parameters in the URI.

Question: You're building an MCP client. What are the two main components you need?

Answer: An MCP Client class and a Client Session.

Question: Your MCP client needs to find out what tools a server offers. What message type should it send?

Answer: ListToolsRequest.

Claude Code and agents (short)

Claude Code is meant to act like another engineer: read the project, plan, then change files. Agents are a language model with tools run in a loop until a goal is met. They lean on environment inspection through tools more than huge static prompts. Parallel Claude Code work uses git worktrees so instances do not stomp the same files. Computer use is tool use over screenshots and UI actions in a controlled environment.

Practice questions

Question: What defines an agent in this course framing?

Answer: A language model with tool access executed repeatedly until a goal is achieved.

Question: What does latency measure in AI systems?

Answer: The time delay between request and response.

Question: Which Claude model is best when maximum speed matters?

Answer: Haiku (course option: Claude 3.5 Haiku).

Question: What are the three main criteria for selecting an AI model?

Answer: Capabilities, speed, and cost.

Closing

For Architect Foundations prep, the durable habits are: measure prompts before you polish them, keep multi-turn history yourself, put structure in tags and schemas, retrieve only what each question needs, and use MCP when you would otherwise invent the same tool glue again. Keep the quiz answers in a flashcard deck and re-check platform docs before you trust a model ID in production.

Comments