Prompt Details
Model
(claude-5-sonnet)
Token size
2,202
Example input
[FAILURE_TRACE]: Query: “What is our refund window for enterprise customers?” Bad answer: “Enterprise customers may request a refund within 14 days of purchase.” Retrieved context shown in trace: Candidate 1, Consumer Refund Policy: “Individual subscriptions have a 14-day refund window.” Candidate 2, refund FAQ repeating the individual-plan rule. Candidate 7, Enterprise MSA Guide: “Enterprise refunds follow the signed MSA and vary by negotiated contract.” Only the first four candidates were sent to generation.
[EXPECTED_BEHAVIOR]: The answer should state that enterprise refund terms are governed by the signed MSA, avoid giving the consumer 14-day window, and cite the enterprise source. Success means the enterprise source enters generation context and the answer preserves the scope distinction.
Example output
### BLOCK 0: INPUT STATUS
Mode: LIVE RUN. Query, bad answer, candidates, cutoff, and expected behavior are captured. The live scoring formula and corpus contents are not supplied.
### BLOCK 1: NORMALIZED INCIDENT
- Query scope: enterprise customers — OBSERVED.
- Bad claim: enterprise refunds are available within 14 days — OBSERVED.
- Supporting source for that wording: consumer policy limited to individual subscriptions — OBSERVED.
- Correct enterprise source: Candidate 7, below the first-four generation cutoff — OBSERVED.
### BLOCK 2: SENTENCE-LEVEL SOURCE TRACE
| Claim | Supporting chunk | Trace status | Contradiction | Evidence state |
|---|---|---|---|---|
| Enterprise customers may request a refund within 14 days | Consumer Refund Policy | SUPPORTED for consumer scope; invalidly transferred | Enterprise MSA Guide says terms vary by contract | OBSERVED |
### BLOCK 3: POOL vs ORDERING VERDICT
Primary stage: ORDERING_LOSS with WRONG_CHUNK_SELECTED.
The correct enterprise source entered the recorded candidates but remained below the generation cutoff. This is not a pool miss. The consumer chunk then won despite a scope mismatch.
### BLOCK 4: ATOMIC CAUSE MAP
| Cause | Failure class | Evidence signal | Evidence state | Explained content |
|---|---|---|---|---|
| Enterprise source ranked below cutoff | ORDERING_LOSS | Candidate 7 vs first-four cutoff | OBSERVED | correct rule absent from generation context |
| Consumer source outranked scope-matched source | WRONG_CHUNK_SELECTED | Candidate 1 is individual-only | OBSERVED | 14-day claim transferred to enterprise |
### BLOCK 5: COUNTERFACTUAL RESULTS
| Cause | Mechanism test | Removal test | Competing cause | Result |
|---|---|---|---|---|
| Enterprise source below cutoff | explains missing authoritative context | moving it inside context makes correct answer plausible | consumer dominance remains possible | SUPPORTED_CAUSE |
| Consumer source wins scope match | directly explains wrong number and scope | suppressing mismatch removes supplied support for bad claim | ordering loss overlaps but is distinct | SUPPORTED_CAUSE |
### BLOCK 6: EVIDENCE-RANKED DIAGNOSIS
Primary cause: scope-insensitive ordering selected consumer evidence over an enterprise source — OBSERVED.
Competing/linked cause: the enterprise source fell below the generation cutoff — OBSERVED.
Confidence state: DIAGNOSIS_SUPPORTED.
### BLOCK 7: MINIMAL REPAIR
Current mechanism → fused candidates ordered without customer-scope precedence.
Proposed change → add enterprise/consumer scope metadata as a deterministic final-ordering feature before cutoff.
Failure class → WRONG_CHUNK_SELECTED / ORDERING_LOSS.
Trade-off → requires reliable scope metadata.
Expected effect → PROJECTED; the enterprise source should enter generation context for enterprise queries, not measured until replayed.
Deferred → embedding replacement and re-chunking; neither targets the observed stage.
### BLOCK 8: REPLAY SPECIFICATION
Replay the exact query. Capture complete candidates, scope metadata, pre/post ordering, generation cutoff, final context, answer, and citation. Confirm if the enterprise source moves inside the cutoff and the consumer rule is not applied. Falsify the repair if the enterprise source remains absent or the answer still copies the consumer rule despite correct context.
### BLOCK 9: ANALYSIS LIMITS
The live pipeline was not replayed. This analysis does not establish general retrieval quality or production causality beyond the supplied incident.
By purchasing this prompt, you agree to our terms of service
CLAUDE-5-SONNET
Take one bad RAG answer apart without turning it into a vague system audit. This prompt traces every claim to supplied context, locates the failure across coverage, pool, ordering, context competition, or synthesis, breaks the incident into single-mechanism causes, challenges each counterfactually, and produces one minimal repair plus an exact replay specification.
...more
Added 1 week ago
