Prompt Details
Model
(claude-4-8-opus)
Token size
833
Example input
[ARGUMENT]: Our company should switch the whole engineering team to a 4-day workweek
[CLAIMS]: Shorter weeks boost productivity; happier staff quit less; we'll attract better hires
[EVIDENCE]: A 2022 UK 4-day-week trial; a Buffer blog post; our own last two sprint velocities being our highest ever
[AUDIENCE]: Company executive leadership team
[PURPOSE]: Approve a permanent policy change
Example output
π§ AUDIT EXECUTION
π SECTION 1 β Executive Assessment
Main Claim: Engineering should move permanently to a 4-day workweek.
Reasoning: Appeal to precedent (UK trial), appeal to authority (Buffer), and inductive generalization from two internal data points.
Strength: Weak β the load-bearing internal evidence is two sprints, far too small to generalize.
Summary: The argument leans on the UK trial and a Buffer post as if they transfer directly to this company, an attribution error β neither studied this team, stack, or client-SLA obligations. It conflates correlation (two high-velocity sprints) with causation (the 4-day model), when those sprints predate any 4-day change. The productivity and retention claims are stacked as if independent, but all three rest on the same thin base. Cherry-picking "our highest ever" two sprints while ignoring the rest of the year is a textbook selective-evidence move.
ποΈ SECTION 2 β Argument Structure
P1: A 2022 UK trial showed 4-day weeks maintained or raised output.
P2: Buffer reports staff are happier on shorter weeks.
P3: Our last two sprints were our highest-velocity ever.
IC-A: Therefore a 4-day week raises our productivity (from P1, P3).
IC-B: Therefore it will cut attrition and attract hires (from P2).
Final: Therefore we should adopt a permanent 4-day week.
P1 (UK trial) ββ
βββΊ IC-A (raises OUR productivity) ββ
P3 (2 sprints)ββ β
βββΊ FINAL (adopt permanently)
P2 (Buffer) βββββββΊ IC-B (retention + hiring) ββββββ
π SECTION 3 β Evidence EVALUATION
UK 2022 4-day-week trial π‘
Source: External pilot study, ~61 UK firms.
β
Real, sizeable, peer-discussed dataset.
β Different companies, roles, and countries β external validity to this team is asserted, not shown. π¨ hasty generalization.
Rating: Moderate strength, but narrow scope β supports "possible," not "will work here."
Buffer blog post π΄
Source: Company marketing/culture blog.
β Self-interested publisher, no methodology, single firm. π¨ appeal to authority.
Rating: Low strength, anecdotal scope.
Two internal sprint velocities π΄
Source: Own recent sprint metrics.
β n=2, and they occurred before the 4-day change β cannot evidence a change not yet made. π¨ post hoc / cum hoc reasoning.
Rating: High severity flaw, decision-critical scope.
π SECTION 4 β Hidden Assumptions
A1 β Transferability. Unstated: results from UK firms apply to this exact team. Reality: sector, client contracts, and time-zone coverage may forbid a dropped day.
A2 β Velocity = value. Unstated: sprint velocity measures real productivity. Reality: velocity is a planning proxy, easily inflated and team-relative.
A3 β Ceteris paribus. Unstated: nothing else caused the two good sprints. Reality: new hires, easier tickets, or crunch could explain them.
β οΈ SECTION 5 β Logical Fallacies
Post hoc / cum hoc β CRITICAL. Error: two strong sprints are credited to a policy not yet in place. Reality: the cause is chronologically impossible. Impact: the central "it works for us" claim collapses.
Hasty generalization β MAJOR. Error: n=2 sprints and other firms' trials generalized to permanent policy. Impact: overstated confidence.
Appeal to authority β MODERATE. Error: Buffer cited as proof. Impact: decorative, not evidential.
π§ SECTION 6 β Reasoning Quality
Coherence β
β the pieces connect logically.
Consistency β
β no internal contradiction.
Validity β οΈ β conclusion doesn't follow with the certainty claimed.
Soundness β β a key premise (sprints prove the model) is false.
Explanatory Power π‘ β plausible story, but rival explanations un-excluded.
βοΈ SECTION 7 β Counterargument Analysis
Coverage gaps during the off-day π΄ β unaddressed; client SLAs may require 5-day support.
Compression, not reduction π΄ β unaddressed; same work in fewer days can raise burnout, the opposite of the retention claim.
Trial β permanent π‘ β partially addressed; novelty effects in trials often fade.
π SECTION 8 β Improvements
Current: "Two sprints were our best ever, so 4 days works." Improved: Run a 90-day internal 4-day pilot, pre-register velocity + defect-rate + on-call-coverage metrics, and compare against the trailing 6-sprint baseline (not the best 2).
Current: cites Buffer as proof. Improved: replace with the UK trial's quantified retention delta (e.g., report the actual % attrition change) and test whether it holds for engineering specifically.
Current: ignores coverage. Improved: add a staggered-day-off model so the team still covers 5 days, and state the target SLA response time it must preserve.
π SECTION 9 β Scorecard
CriterionScore /10NotesLogical Consistency7Internally coherentEvidence Strength3Thin, mis-timed, self-interested sourcesAssumption Transparency3Key assumptions buriedFallacy Resistance2Central post-hoc errorConclusion Validity3Overreaches the data
Overall Score: 36 / 100 β Verdict: WEAKLY SUPPORTED
π§Ύ FINAL SUMMARY
Strengths: clear structure; a genuinely relevant external trial; a testable proposal.
Critical Weaknesses: central evidence is chronologically impossible; sample sizes tiny; coverage and burnout counterarguments ignored.
Core Problem: The argument proves that a 4-day week is worth piloting; it does not prove it will work for this team permanently. It borrows others' results and its own pre-change metrics to skip the one test that matters.
Deeper Issue: The root error is treating a proxy metric measured before the intervention as evidence for the intervention β mistaking a pre-existing trend for a projected effect.
β
DEMONSTRATES
π’ Premise/conclusion decomposition and mapping
π’ Source credibility and timing analysis
π’ Hidden-assumption surfacing
π’ Named-fallacy detection with severity
π’ Counterargument stress-testing
π’ Quantified, targeted remediation
By purchasing this prompt, you agree to our terms of service
CLAUDE-4-8-OPUS
Stress-test any argument like a debate coach. Feed in a claim, its supporting points, the evidence cited, your audience, and the decision at stake β get back a rigorous 9-section audit: premise map, evidence traffic-lights, hidden assumptions, named fallacies with severity, a counterargument sweep, targeted fixes, and a 100-point scorecard with verdict. For analysts, students, writers, lawyers, and anyone deciding under scrutiny
...more
Added 1 week ago
