A Claude-based research assistant begins producing responses that confidently contradict its retrieved source documents despite no change to the retrieval pipeline.
Which two diagnostic actions most directly identify the root cause of this behavior? (Select two.)
A. Determine whether the failure reproduces on the previous model version to test for a model mismatch.
B. Reduce the temperature setting to lower response variance across all query types.
C. Increase the context-window size to allow more retrieved chunks per query.
D. Inspect the system-prompt grounding instructions to determine whether citation constraints remain intact.
E. Switch the retrieval index to a denser embedding model to improve chunk-relevance scores.
正解:A,D
解説: (Pass4Test メンバーにのみ表示されます)
質問 2:
You are running a controlled experiment to compare two prompts and must complete the design steps before executing the experiment.
Which two steps must be completed BEFORE running the experiment with random assignment? (Select two.) Each correct answer presents part of the solution.
A. Decide whether to promote, reject, or iterate the candidate based on the analysis.
B. Analyze the results against the predefined success metric and significance threshold.
C. Determine the minimum detectable effect size and the sample size needed for power.
D. Document the recommendation, the trade-offs accepted, and the alternatives considered.
E. Define the hypothesis and the primary success metric for the comparison.
正解:C,E
解説: (Pass4Test メンバーにのみ表示されます)
質問 3:
You are integrating human review into a high-volume classification pipeline where reviewing every output is infeasible.
Which sampling strategy best balances throughput with quality oversight?
A. Inverse sampling that reviews only high-confidence routine outputs and skips low-confidence and high- impact outputs.
B. No sampling, relying entirely on user complaints to reveal quality and safety problems after they affect users.
C. Risk-stratified sampling that reviews all low-confidence and high-impact outputs and a smaller random sample of high-confidence routine outputs.
D. Universal review of every output regardless of confidence or throughput impact.
正解:C
解説: (Pass4Test メンバーにのみ表示されます)
質問 4:
A loan pre-qualification assistant shows 94 percent approval recommendations that match the human underwriter decision. The fairness team has reviewed approval rate parity across protected groups and reported no significant difference. A board member has asked whether this evidence is sufficient to declare the assistant fair.
Which two Discernment-competency findings should you report? (Select two.) Each correct answer presents part of the solution.
A. Match with human underwriters does not establish freedom from underwriter-introduced bias.
B. A larger sample is needed before any meaningful fairness claim can be made about the model.
C. Approval rate parity does not by itself assess error rate parity across protected groups.
D. The fairness team's review process likely missed at least some of the protected groups studied.
E. The 94 percent match rate is sufficient evidence of fairness for the assistant's decisions.
正解:A,C
解説: (Pass4Test メンバーにのみ表示されます)
質問 5:
You are preparing an operational runbook for a Claude-based service.
Which content is essential to include in the runbook?
A. Escalation paths and rollback procedures only, without alert definitions, triage steps, or dashboard references to guide initial incident response.
B. Dashboard and log references only, without alert definitions, triage steps, escalation paths, or rollback procedures for the on-call engineer to act on.
C. Common alerts and their triage steps, escalation paths, rollback procedures, and references to the relevant dashboards and logs.
D. Alert definitions and triage steps only, without escalation paths, rollback procedures, or references to dashboards and logs for on-call use.
正解:C
解説: (Pass4Test メンバーにのみ表示されます)
質問 6:
You are evaluating an evaluation set used to score a Claude-based hiring-support tool. The set is drawn from one geographic region and one tenure band.
Which response is most appropriate?
A. Reduce the evaluation set further to a single demographic subgroup to simplify score interpretation, narrowing coverage rather than expanding it to match the intended user population.
B. Continue using the narrow evaluation set drawn from one region and one tenure band because the existing benchmark scores are already high on that subset.
C. Expand the evaluation set to cover the geographic regions and tenure bands the tool will serve, and rescore the system on the expanded set before broader release.
D. Discard all quantitative evaluation and replace it with qualitative impressions collected from a small, convenience-selected group that may not represent the tool's full user population.
正解:C
解説: (Pass4Test メンバーにのみ表示されます)




0 お客様のコメント