Limited-Time Offer: Enjoy 50% Savings! Ends in 00h 00m 00s Coupon code: 50OFF
Skip to content

Free Anthropic Claude Certified Architect - Foundations CCAR-F Exam Questions

Page: 1 / 11 Total 152 questions

Want more questions? Get Premium Access.

Question 1

You are building a structured data extraction system using Claude. The system extracts information from unstructured documents, validates the output using JavaScript Object Notation (JSON) schemas, and maintains high accuracy. It must handle edge cases gracefully and integrate with downstream systems.

Your system has been running for 3 weeks and human reviewers have corrected 847 extractions. Analysis reveals a recurring pattern: when recipes use informal measurements like ''a handful'' or ''a splash,'' the model either invents specific amounts or leaves fields empty---accounting for 23% of all corrections.

How should you use this feedback to improve extraction accuracy?

Correct Answer: B. Add few-shot examples to your prompt demonstrating correct handling of informal measurements---extracting them verbatim rather than converting or omitting them.
Explanation:

The reviewer corrections have exposed a narrow and repeatable interpretation failure. The desired policy is clear: informal measurements are valid source values and must be preserved verbatim rather than normalized into invented quantities or treated as missing. This behavior can be communicated efficiently through targeted few-shot examples.

Anthropic recommends examples for demonstrating expected behavior and improving consistency. Examples can pair source phrases such as ''a handful of spinach,'' ''a splash of vinegar,'' and ''a pinch of salt'' with outputs that retain handful, splash, and pinch exactly. Additional counterexamples can show that Claude must not convert these phrases into grams, millilitres, or estimated serving quantities. (https://docs.anthropic.com/en/docs/about-claude/use-case-guides/ticket-routing)

Option A introduces a substantially heavier training workflow for a problem that can be addressed directly through the prompt. Option C creates a parallel extraction mechanism based on pattern matching; it will be brittle across linguistic variations and may populate a value without understanding its relationship to the correct ingredient. Option D adds useful classification metadata, but it does not instruct Claude to preserve the original measurement instead of inventing or omitting it.

The revised prompt should combine an explicit verbatim-extraction rule with several varied examples derived from the corrected cases, followed by regression evaluation against the identified failure set.

Official references/topics: Few-Shot Examples, Feedback-Driven Prompt Improvement, Verbatim Extraction, Regression Evaluation.


Question 2

You are building a multi-agent research system using the Claude Agent SDK. A coordinator agent delegates to specialized subagents: one searches the web, one analyzes documents, one synthesizes findings, and one generates reports. The system researches topics and produces comprehensive, cited reports.

The synthesis agent completes its initial pass but flags that three key research questions remain unanswered because the web-search and document-analysis agents did not find relevant information on those specific subtopics. The coordinator currently proceeds directly to report generation, producing reports with incomplete coverage.

What change would most effectively improve research completeness?

Correct Answer: B. Have the coordinator evaluate the synthesis output for gaps, then re-delegate to web search and document analysis with targeted queries before invoking synthesis again.
Explanation:

Option B introduces an evaluator-and-refinement loop at the correct orchestration layer. The coordinator already owns the research plan and delegation decisions, so it should inspect the synthesis result against the required questions, identify coverage gaps, and issue focused follow-up assignments. Anthropic's description of its multi-agent research system follows this pattern: the lead agent synthesizes returned findings, determines whether additional research is required, and creates new subagents or refines its strategy before producing the final result. Increasing the initial query breadth, option A, may generate additional irrelevant material and cannot guarantee that unforeseen gaps will be covered. Option C merely documents the incompleteness instead of correcting it. Option D weakens role separation by giving the synthesis agent search capabilities, increasing tool complexity and bypassing the coordinator's centralized tracking. Targeted re-delegation preserves specialized responsibilities and creates an observable sequence of research, evaluation, refinement, and resynthesis. The coordinator should also maintain explicit coverage criteria and limit the number of refinement rounds so the system improves completeness without entering an uncontrolled research loop.


Question 3

You are building developer-productivity tools using the Claude Agent SDK. The agent helps engineers explore unfamiliar codebases, understand legacy systems, generate boilerplate code, and automate repetitive tasks. It uses the built-in tools---Read, Write, Bash, Grep, and Glob---and integrates with Model Context Protocol (MCP) servers.

Engineers frequently ask the agent to cross-reference code changes with Jira tickets during reviews---checking ticket descriptions, acceptance criteria, and recent comments. This currently requires manually copying and pasting content into conversations. The team wants the agent to access this standard Jira ticket data directly.

What is the most effective approach?

Correct Answer: D. Integrate an existing Jira MCP server that exposes tickets, comments, and metadata through discoverable tool interfaces.
Explanation:

Option D uses the established integration mechanism without creating unnecessary infrastructure. Anthropic's Claude Code MCP documentation specifically recommends connecting an MCP server when users repeatedly copy information from an external system, such as an issue tracker, into conversations. Once connected, the server exposes Jira operations through named, schema-defined tools that Claude can discover and invoke directly. This allows the agent to retrieve ticket descriptions, acceptance criteria, comments, and metadata while preserving the server's authentication and access controls.

Option A exposes authentication details to shell commands and requires the agent to construct requests and parse responses repeatedly. Option B is justified only when no existing server provides the required capabilities or the workflow requires highly specialized operations. The question requires standard Jira information, so building and maintaining another server is unnecessary. Option C creates stale, manually synchronized copies that lose permissions, current comments, and authoritative ticket state. The team should therefore select a trusted Jira MCP implementation, configure appropriate read-only scopes where possible, and verify its tool descriptions and data-access boundaries before enabling it.


Question 4

You are using Claude Code to accelerate software development. Your team uses it for code generation, refactoring, debugging, and documentation. You need to integrate it into your development workflow with custom slash commands, CLAUDE.md configurations, and understand when to use plan mode vs direct execution.

You need to add a date validation check ensuring event dates are in the future. This requires adding a conditional statement to one existing function in a single file.

What is the most appropriate approach?

Correct Answer: A. Use direct execution to make the change.
Explanation:

This change is narrow, localized, and already defined: add one conditional validation check to an existing function in a single file. A separate planning phase would introduce process overhead without resolving meaningful architectural uncertainty. Direct execution allows Claude to read the function, implement the condition, and run the relevant focused tests.

Anthropic explicitly states that plan mode adds overhead and should generally be skipped when the scope is clear and the fix is small. Planning is most valuable when the approach is uncertain, multiple files are affected, or the code is unfamiliar. Anthropic's practical rule is that when the required diff can be described in one sentence, direct implementation is appropriate. (https://code.claude.com/docs/en/best-practices)

Option B allocates unnecessary reasoning effort to straightforward validation logic. Options C and D exaggerate the complexity of a single-function change. Broader impact analysis would be justified only if the requirement altered reservation semantics, time-zone rules, persistence behavior, or public interfaces---none of which is stated.

The implementation should still include verification. Claude should add or update tests for a future date, the current date, and a past date, then run the narrowest relevant test command. Direct execution does not mean unverified execution.

Official references/topics: Direct Execution; Plan-Mode Selection; Small Scoped Changes; Focused Verification.


Question 5

When researching ''renewable-energy adoption,'' the web-search agent returns recent statistics showing 35% adoption in 2024, while the document-analysis agent extracts an 18% adoption figure from an internal 2021 report. The synthesis agent incorrectly treats the figures as contradictory instead of recognizing that they may show growth over time. What change would best enable the synthesis agent to interpret such temporal differences correctly?

Correct Answer: A. Require subagents to include publication dates and data-collection periods in their structured outputs.
Explanation:

The two statistics cannot be compared correctly without temporal metadata. Option A preserves when each value was measured, not merely when its document was retrieved. The structured record should ideally distinguish publication date, observation period, geographic scope, population, methodology, unit, and metric definition. The synthesis agent can then determine whether the figures represent a trend, conflicting measurements of the same period, or incomparable populations.

Anthropic's multi-agent research guidance emphasizes specifying output formats and source requirements so downstream agents receive the context required for accurate synthesis. Its structured-output capability enables these provenance fields to be required and schema validated.

Option B removes valuable historical evidence and still does not guarantee that returned sources describe the same measurement period. Option C assumes newer automatically means more reliable and could discard an authoritative historical baseline. Option D preserves older material but imposes an interpretation without examining whether methodology, geography, or definitions changed. Explicit temporal and methodological metadata allows the synthesis agent to state a defensible conclusion such as apparent growth from 18% in 2021 to 35% in 2024 while preserving any comparability caveats.


Question 6

You are building a structured data extraction system using Claude. The system extracts information from unstructured documents, validates the output using JavaScript Object Notation (JSON) schemas, and maintains high accuracy. It must handle edge cases gracefully and integrate with downstream systems.

Your extraction pipeline processes invoices and extracts line items, subtotals, tax amounts, and grand totals. During evaluation, you discover that in 18% of extractions, the sum of extracted line item amounts doesn't match the extracted grand total---sometimes due to OCR errors in the source document, sometimes due to extraction mistakes by the model. Downstream accounting systems reject records with mismatched totals.

What's the most effective approach to improve extraction reliability?

Correct Answer: D. Add a ''calculated_total'' field where the model sums extracted line items alongside a ''stated_total'' field. Flag records for human review when values differ.
Explanation:

The pipeline must preserve source evidence while making inconsistencies explicit. Option D records the amount stated on the invoice separately from the total derived from extracted line items. A mismatch then becomes a machine-detectable validation condition rather than an invisible extraction defect.

This approach is superior because it does not silently overwrite source data or ask another model to guess which value is correct. Anthropic's evaluation guidance recommends automated, code-based grading whenever the criterion can be expressed deterministically. Arithmetic reconciliation is precisely such a criterion. (https://docs.anthropic.com/en/docs/build-with-claude/develop-tests) In production, the summation should preferably be calculated by application code using normalized decimal values, even though the option describes the model populating calculated_total. The essential design principle remains the same: preserve stated_total, compute an independent total, compare them, and route discrepancies for adjudication.

Option A might improve behavior but cannot resolve genuine OCR corruption and could encourage the model to modify extracted values merely to create mathematical consistency. Option B introduces a second probabilistic judgment without new evidence. Option C is unacceptable for accounting data because it fabricates adjusted amounts and destroys fidelity to the invoice.

The schema should therefore expose both values and attach a validation status or review reason when they differ.

Official references/topics: Deterministic Validation; Human-in-the-Loop Review; Structured Output Design; Source-Fidelity Controls.


Question 7

Production reviews reveal inconsistent handling of uncertainty in final reports. Sometimes conflicting subagent findings are synthesized into a single confident statement, losing important nuance, while other reports over-hedge with excessive qualifications and become unhelpful. The web-search agent returns, ''Industry analysts estimate a $50 billion market size, although methodologies vary.'' The document-analysis agent returns, ''A peer-reviewed study estimates $35 billion, with a $7 billion 95% confidence interval.'' The coordinator either selects one estimate arbitrarily or produces a vague $35--$50 billion range. What systematic approach best addresses this?

Correct Answer: A. Instruct the synthesis agent to structure reports with explicit sections distinguishing well-established findings from contested findings while preserving each source's characterization and methodological context.
Explanation:

Option A preserves disagreement as meaningful evidence. The report should identify the two estimates separately, explain that they arise from different methodologies and source types, retain the peer-reviewed study's stated confidence interval, and classify the market size as contested rather than manufacturing false agreement. It may then explain what additional evidence would resolve the discrepancy.

Anthropic's research prompting guidance recommends developing competing hypotheses, tracking confidence, verifying across sources, and retaining structured research notes. This supports calibrated synthesis rather than arbitrary selection or vague hedging.

Option B discards potentially important minority evidence and incorrectly treats corroboration count as equivalent to reliability. Option C creates numerical precision unsupported by the original sources; confidence scores from different agents are not automatically comparable, and averaging incompatible estimates can be meaningless. Option D introduces survivorship bias by suppressing uncertainty before synthesis. A systematic reporting structure distinguishes consensus, disagreement, evidence quality, methodological limitations, and unresolved questions. That produces an actionable briefing without overstating certainty or weakening every conclusion with generic qualifications.


Question 8

You are building a multi-agent research system using the Claude Agent SDK. A coordinator agent delegates to specialized subagents: one searches the web, one analyzes documents, one synthesizes findings, and one generates reports. The system researches topics and produces comprehensive, cited reports.

A user expands the research system beyond its original web-search agent by adding specialized data sources. A financial API agent returns structured JSON containing revenue, margins, and growth rates. A news-monitoring agent returns prose summaries of recent developments. A patent-analysis agent returns structured lists of technology areas. The synthesis agent combines these results into executive briefings. Currently, it converts everything into bullet points, causing financial comparisons to lose tabular clarity and news summaries to lose their narrative flow.

What change would most improve briefing quality?

Correct Answer: C. Update the synthesis agent to render each content type appropriately---for example, financial data as tables, news as prose, and patent areas as structured lists.
Explanation:

Option C preserves the information structure that makes each source useful. Financial metrics share comparable fields and therefore benefit from rows, columns, aligned units, and reporting periods. News findings require connected prose to preserve chronology and causal relationships, while patent technology areas are naturally represented as categorized lists. Anthropic's output-consistency guidance recommends specifying the exact output format needed for the task rather than relying on an unspecified default. Anthropic's discussion of its multi-agent research system also recognizes specialized output stages for reports, structured data, and visualizations because specialist prompts can produce better results than generic coordinator processing. Option A destroys the comparative structure of numerical data. Option D can provide a useful provenance contract internally but does not determine how the executive briefing should present heterogeneous content. Option B risks creating a lowest-common-denominator representation that discards source-specific advantages. The synthesis contract should preserve normalized facts and provenance internally while directing the report generator to select presentation forms according to the content's semantic structure and the executive reader's needs.


Question 9

You are building developer productivity tools using the Claude Agent SDK. The agent helps engineers explore unfamiliar codebases, understand legacy systems, generate boilerplate code, and automate repetitive tasks. It uses the built-in tools (Read, Write, Bash, Grep, Glob) and integrates with Model Context Protocol (MCP) servers.

1.5An engineer asks the agent to understand how the caching layer works before adding a new cache invalidation trigger. After initial Grep searches, the agent has identified that caching logic spans 15 files including decorators, middleware, and service classes (~6,000 lines total).

What's the most effective next step for building understanding while managing context constraints?

Correct Answer: D. Analyze imports and class hierarchies to identify the base cache class. Read that file to understand the interface, then trace specific invalidation implementations.
Explanation:

The correct objective is to construct an architectural map before consuming the full implementation. Identifying the base cache abstraction, its interface, and the classes that implement or invoke it gives the agent a dependency-guided path through the code. It can then inspect only the invalidation implementations and integration points relevant to the proposed trigger.

This approach protects the context window. Anthropic states that every file read occupies context and that model performance can deteriorate as the window fills. Its Claude Code guidance warns against unbounded investigation that reads large numbers of files and recommends narrowing the exploration or delegating it. (https://code.claude.com/docs/en/best-practices)

Option A is too lexical: searching only for invalidate or expire can miss event-driven invalidation, overridden methods, cache-key mutation, and generic interface calls. Option B loads approximately 6,000 lines without first establishing relevance. Option C assumes that filename patterns and file size correlate with architectural importance; the largest files may contain incidental code while a small interface defines the entire design.

Option D follows control and type relationships rather than arbitrary file order. After reading the base class, the agent can search for subclasses, imports, construction sites, middleware hooks, and calls to the invalidation contract, progressively expanding only where evidence requires it.

Official references/topics: Context-Efficient Exploration; Dependency-Guided Reading; Architectural Interfaces; Narrowly Scoped Investigation.


Question 10

You are building a customer support resolution agent using the Claude Agent SDK. The agent handles high-ambiguity requests like returns, billing disputes, and account issues. It has access to your backend systems through custom Model Context Protocol (MCP) tools (get_customer, lookup_order, process_refund, escalate_to_human). Your target is 80%+ first-contact resolution while knowing when to escalate.

You're implementing the escalation logic for when the agent should call escalate_to_human. Your team proposes four different approaches for triggering escalation.

Which approach will most reliably identify cases that genuinely require human intervention?

Correct Answer: B. Instruct the agent to escalate when the customer requests a human, when the issue requires policy exceptions, or when the agent cannot make meaningful progress.
Explanation:

Option B identifies escalation through direct operational criteria rather than indirect proxies. A customer's explicit request for a person is unambiguous. A required policy exception defines an authorization boundary. An inability to make meaningful progress identifies a case where continued autonomous execution is no longer productive.

Anthropic describes agents as systems that act, observe results, adjust, and continue until the task is completed or human input is required. Its trustworthy-agent guidance also emphasizes that agents must recognize when to pause rather than pushing through uncertainty or decisions that only a human can settle. (https://www.anthropic.com/research/trustworthy-agents)

Option A is too rigid for high-ambiguity support cases. Maintaining exhaustive mappings for every issue, product, and customer segment recreates the brittleness that agentic reasoning is intended to avoid. Option C treats repeated tool use as a substitute for semantic progress; one definitive authorization failure may require immediate escalation, while several legitimate diagnostic calls may not. Option D confuses emotional tone with operational necessity. A calm customer may require a policy exception, while a frustrated customer may still have a straightforward automated resolution.

The escalation criteria should be encoded in the system instructions and supported by evaluations covering explicit requests, authorization limits, unresolved ambiguity, repeated non-progress, and successful self-service cases.

Official references/topics: Human-control checkpoints, escalation criteria, authorization boundaries, progress-aware agents.


Question 11

Your automated review calls the Claude API for each pull request, using tool_use with a report_findings tool that returns a JSON array of finding objects. Each object contains file_path, line_number, severity, category, and description. During testing on a large pull request touching more than 30 files, the response reaches the max_tokens limit and is truncated in the middle of the JSON, causing your pipeline's parser to fail. What is the most effective way to handle this?

Correct Answer: A. Split the review into multiple API calls that each analyze a subset of the changed files, and then merge the resulting findings arrays.
Explanation:

Option A reduces the maximum output required from any single response while preserving the structured schema and complete severity range. The pipeline can partition files into coherent groups, execute bounded reviews, validate each returned array, and merge and deduplicate findings using stable fields such as file path, line number, category, and description.

Anthropic's stop-reason documentation confirms that max_tokens means generation reached the configured output limit and the response must be treated as incomplete. Structured output constraints can guarantee schema-valid generation when completion succeeds, but they cannot create unlimited output capacity. A large findings array can still exceed the available token budget.

Option B may postpone the failure but provides no durable guarantee for still-larger pull requests, and aggressively shortening descriptions may eliminate necessary evidence. Option C abandons machine-validated structure without reducing the amount of generated content. Option D deliberately suppresses medium- or low-severity findings and repeats an oversized request rather than addressing its scope. Partitioning establishes predictable output bounds, supports targeted retries, retains every required finding category, and prevents a single truncated response from invalidating the complete review.


Question 12

You are building developer-productivity tools using the Claude Agent SDK. The agent helps engineers explore unfamiliar codebases, understand legacy systems, generate boilerplate code, and automate repetitive tasks. It uses the built-in tools---Read, Write, Bash, Grep, and Glob---and integrates with Model Context Protocol (MCP) servers.

After adding an MCP server with specialized code-refactoring tools---extract_function, rename_variable, and inline_function---you notice that the agent still uses basic text manipulation through Write and Bash sed commands for refactoring tasks. The MCP server is connected and healthy. Examining the configuration, you find that each MCP tool has a minimal description such as, ''extract_function: Extracts a function from code.''

What is the most effective way to improve adoption of the MCP refactoring tools?

Correct Answer: C. Enhance the MCP tool descriptions to explain when each tool is preferable to text manipulation and clarify expected inputs and outputs.
Explanation:

Option C corrects the weak selection signal presented to the model. Claude chooses among available tools using their names, descriptions, parameter schemas, and the current request. ''Extracts a function from code'' does not explain whether the tool understands syntax trees, preserves imports, updates call sites, validates scope, or offers advantages over Write and sed. Anthropic identifies prompt-engineering tool descriptions as one of the most effective ways to improve agent tool use. Descriptions should state what the operation performs, when it should be selected, what inputs are required, what output it returns, and any limitations.

Option A adds a separate probabilistic routing layer without improving the tool contract Claude ultimately sees. Option B ignores the server's intended value. Option D removes a broadly useful capability and may prevent unrelated edits without guaranteeing that the MCP tools are used correctly. Each refactoring tool should instead describe its semantic behavior and contrast it with plain text manipulation---for example, that rename_variable performs scope-aware symbol renaming and updates references. Clear schemas, concrete examples, and evaluation against real refactoring tasks should accompany the improved descriptions.


Question 13

You are building a structured data extraction system using Claude. The system extracts information from unstructured documents, validates the output using JavaScript Object Notation (JSON) schemas, and maintains high accuracy. It must handle edge cases gracefully and integrate with downstream systems.

Your system must extract event details from calendar invitations and output JSON that strictly conforms to a schema with fields for title, date, time, location, and attendees. Downstream systems reject any malformed or non-conformant JSON.

What approach provides the most reliable schema compliance?

Correct Answer: C. Define a tool with your target schema as input parameters and have Claude call it with the extracted data.
Explanation:

A tool definition converts the desired extraction structure into an explicit machine-readable contract. Claude returns the event information inside a tool_use block, with the tool arguments corresponding to the properties defined by the tool's input_schema. Anthropic specifies that custom tool parameters are described using JSON Schema, allowing the application to extract the structured arguments directly rather than attempting to recover JSON from ordinary prose. For current implementations, adding strict: true to the tool definition provides guaranteed conformance of tool-call inputs to the declared schema. (https://docs.anthropic.com/en/docs/agents-and-tools/tool-use/implement-tool-use)

Options A, B, and D remain prompt-based formatting techniques. They may improve the probability of valid JSON, but none creates the same schema-enforced interface. Prefilling an opening brace constrains the beginning of the response without guaranteeing valid field names, required properties, or data types. Retry logic detects failures only after generation and adds latency. Detailed formatting instructions can still produce malformed or structurally incorrect output.

Anthropic now also provides Structured Outputs for direct, schema-validated JSON responses. Within the options presented, however, a schema-defined tool is the only approach that establishes an explicit structured-output boundary rather than relying primarily on text-generation compliance. (https://docs.anthropic.com/en/docs/test-and-evaluate/strengthen-guardrails/increase-consistency)

Official references/topics: Tool Definitions, JSON Schema Input Contracts, Strict Tool Use, Structured Outputs.


Question 14

You are integrating Claude Code into your Continuous Integration/Continuous Deployment (CI/CD) pipeline. The system runs automated code reviews, generates test cases, and provides feedback on pull requests. You need to design prompts that provide actionable feedback and minimize false positives.

Your CI pipeline performs security-focused code reviews on approximately 50 pull requests daily, currently costing $150 per day using the synchronous API. Reviews are non-blocking---developers merge after tests pass and address findings in follow-up commits. You are evaluating the Message Batches API for its 50% cost reduction.

What factor most determines whether batch processing is appropriate for this use case?

Correct Answer: C. Whether review feedback arriving up to 24 hours after pull-request creation remains actionable.
Explanation:

Option C identifies the fundamental trade-off introduced by batch processing: lower cost in exchange for asynchronous completion and potentially substantial latency. Anthropic states that most Message Batches complete within one hour, but results may become available only when all requests finish or after 24 hours, whichever occurs first. Therefore, the workflow must remain useful even if security findings arrive considerably later than they would through synchronous requests.

The reviews are explicitly non-blocking, so batch processing can be suitable if developers can still act on delayed findings through follow-up commits. Option B describes an implementation requirement because batch results may be returned out of submission order, but Anthropic provides a deterministic solution through each request's unique custom_id. Option A incorrectly focuses on near-instant feedback, which batch processing does not provide. Option D is not decisive because batches support independent Messages API requests containing system prompts, tool use, and multi-turn conversation content. The primary decision is therefore whether the maximum practical feedback delay is acceptable to the development workflow. Anthropic Message Batches documentation


Question 15

You are building a multi-agent research system using the Claude Agent SDK. A coordinator agent delegates to specialized subagents: one searches the web, one analyzes documents, one synthesizes findings, and one generates reports. The system researches topics and produces comprehensive, cited reports.

Production reviews reveal inconsistent handling of uncertainty in final reports. Sometimes conflicting subagent findings are synthesized into a single confident statement, losing important nuance, while other reports use excessive qualifications and become unhelpful. The web-search agent returns, ''Industry analysts estimate a $50 billion market size, although methodologies vary.'' The document-analysis agent returns, ''A peer-reviewed study estimates $35 billion, with a $7 billion 95% confidence interval.'' The coordinator either selects one estimate arbitrarily or produces a vague $35--$50 billion range.

What systematic approach best addresses this?

Correct Answer: D. Instruct the synthesis agent to distinguish well-established findings from contested findings explicitly, preserving each source's original uncertainty, methodology, and supporting evidence.
Explanation:

Option D preserves the evidence instead of manufacturing certainty. The two estimates are not directly interchangeable: one is an industry estimate with unspecified methodology, while the other is a peer-reviewed estimate with an explicit confidence interval. Converting both into model-generated confidence scores and calculating a weighted average would create a new figure that neither source reported and that may have no statistical validity. Anthropic's hallucination-reduction guidance recommends making claims auditable through quotations, citations, and supporting evidence rather than presenting unsupported synthesis as fact. Its Citations documentation similarly emphasizes retaining the exact source passages supporting individual claims. Filtering uncertain findings, option B, would remove decision-relevant information. Requiring two-source corroboration, option C, could also discard credible evidence concerning emerging or specialized subjects. The synthesis agent should report the estimates separately, explain their methodological differences, identify which findings are strongly supported or disputed, and state what evidence would resolve the disagreement. This produces calibrated, useful reporting without arbitrary selection, excessive hedging, or false precision.