docs(design): #574 spec A — role-scoped scoring + abstention + per-mode decision contract (design doc) (#584)
- docs(design): #574 spec A — role-scoped scoring + abstention + canonical per-mode decision contract (design doc)
Implements the #574 normative overlay’s P0 gaps 1-2 + C3 + C5 as a design-first document (no code or prompt change ships here), and rules that the E4-named B1 residual (seat-level severity-band anchoring) is folded INTO spec A rather than split to a follow-up issue.
Core settled decisions: contract-carried eligible_roles/owner_role (Schema 13.2, reviewer-conditional); DA scores exactly its remit dimension (D3), findings stay un-scoped; two-form not_assessed abstention with fail-closed [DIMENSION-UNASSESSED]; two-stage panel evaluation (per-dimension quantifiers over eligible seats); per-seat Failure Condition Checks / Editorial Decision retired in favor of executable trigger binding; reject_or_major_revision removed with a pre-committed fatal/repairable block split; exhaustive v2 condition sets with a reachable Minor path; a canonical per-mode decision authority table with a defrift guard; a new per-seat check_phase_conformance.py (plan adequacy, blind-phase content containment, trigger binding, dissent cap, anchor presence).
Co-Authored-By: Claude Fable 5 noreply@anthropic.com Claude-Session: https://claude.ai/code/session_01LWBuqyFwftLpTpBGAJbrho
- docs(design): spec A round-1 review fixes — grammar patterns 7-9, D6 venue-fit dimension, role binding, mandatory manuscript input, DA table-aware anchor gate, abstention enumeration
Closes all 5 P1 + 1 P2 from codex (gpt-5.6-sol xhigh) round 1: the §8.2 condition sets now have matching §7.1 closed-grammar entries; full mode gains the EIC-owned mandatory venue_fit_and_contribution dimension (D6) so an out-of-scope-but-sound paper can no longer mechanically Accept; check_phase_conformance.py gains required –role (dispatch-assigned, must equal contract_role) and required –manuscript (omission exit 2, never a skip) plus a panel-side –roles cross-check; the anchor gate is role-aware for the DA’s committed issue-table grammar; the §15 enumeration covers partial-abstention states (24,576 full-mode profiles).
Co-Authored-By: Claude Fable 5 noreply@anthropic.com Claude-Session: https://claude.ai/code/session_01LWBuqyFwftLpTpBGAJbrho
- docs(design): spec A round-2 review fixes — no fatality without pre-commitment (dissent exception), metadata envelope for the leakage exemption, pairwise trigger distinctness
Closes the 3 round-2 P1s from codex (gpt-5.6-sol xhigh): block_class: fatal is forbidden on a dissented dimension (reject authority never rides an un-pre-committed judgment; the concern lands as a repairable block + finding); check_phase_conformance.py gains required –metadata so the leakage exemption is defined purely over checker inputs (no manuscript-side title extraction); Phase-1 trigger distinctness is specified as pairwise over the normalized three-set with all three collision fixtures in the test plan.
Co-Authored-By: Claude Fable 5 noreply@anthropic.com Claude-Session: https://claude.ai/code/session_01LWBuqyFwftLpTpBGAJbrho
- docs(design): spec A round-3 review fix — DA-CRITICAL terminal consistency gate
Closes the round-3 P1 from codex (gpt-5.6-sol xhigh): role scoping had removed the mechanical expressibility of the ‘validated/unresolved DA-CRITICAL blocks Accept’ iron rule (in v1 the DA’s own dimension block carried it into F1; in v2 the DA scores only D3). New §8.3: a mechanical accept coexisting with a VALIDATED or UNRESOLVED DA-CRITICAL adjudication emits [DA-CRITICAL-VS-ACCEPT] and escalates to the user instead of silently finalizing — an escalation gate in the #518 review-trigger shape, never a decision change; REJECTED-with-rationale does not trigger it (#574 B1). Checker verification, delivery-surface row, decision 14, and gate fixtures added.
Co-Authored-By: Claude Fable 5 noreply@anthropic.com Claude-Session: https://claude.ai/code/session_01LWBuqyFwftLpTpBGAJbrho
- docs(design): spec A round-4 review fixes — machine-addressable DA adjudication contract, mandatory-only fatality, pinned Phase-1 line grammar
Closes the 3 round-4 P1s from codex (gpt-5.6-sol xhigh): §8.3 gains a pinned adjudication grammar (DA CRITICAL rows carry stable IDs C1..Cn; sprint synthesis always emits da_critical_adjudications with total 1:1 ID matching, per-REJECTED rationale lines, and a count-checked marker; DA-table parsing is a shared helper across both checkers — correcting the prior false already-parses claim); block_class/what_triggers_fatal become mandatory-dimension-only (a fatal D5 block feeding Minor Revision asserted unfixable and fix-in-minor simultaneously; fatal atoms over non-mandatory scope are CONTRACT-INVALID as unsatisfiable-by-construction fail-open); §6.1 pins the Phase-1 scoring-plan line grammar (anchored full lines, exactly-once, single-line values) so trigger binding parses one canonical form. Enumeration recomputed (13,824 full / 12 mf); fixtures extended.
Co-Authored-By: Claude Fable 5 noreply@anthropic.com Claude-Session: https://claude.ai/code/session_01LWBuqyFwftLpTpBGAJbrho
Co-authored-by: Claude Fable 5 noreply@anthropic.com
版权所有:中国计算机学会技术支持:开源发展技术委员会
京ICP备13000930号-9
京公网安备 11010802047560号
Academic Research Skills for Claude Code
简体中文版 | 繁體中文版 | 日本語版 | 한국어
A comprehensive suite of Claude Code skills for academic research, covering the full pipeline from research to publication.
Install in 30 seconds (Claude Code CLI / VS Code / JetBrains, v3.7.0+):
Then try
/ars-planto walk through your paper structure via Socratic dialogue, or jump to Quick install for prerequisites and the traditional symlink flow.Why human-in-the-loop, not full automation?
Lu et al. (2026, Nature 651:914-919) built The AI Scientist — the first fully autonomous AI research system to publish a paper through blind peer review at a top-tier ML venue (ICLR 2025 workshop, score 6.33/10 vs workshop average 4.87). Their Limitations section enumerates the failure modes that any fully-autonomous AI research pipeline inherits: implementation bugs, hallucinated results, shortcut reliance, bug-as-insight reframing, methodology fabrication, frame-lock, citation hallucinations.
ARS is built on the premise that a human researcher augmented by AI avoids these failure modes better than either alone. Stage 2.5 and Stage 4.5 integrity gates run a 7-mode blocking checklist (see
academic-pipeline/references/ai_research_failure_modes.md); the reviewer offers an opt-in calibration mode that measures its own FNR/FPR against a user-supplied gold set.Zhao et al. (2026-05) audited 111M references across 2.5M papers on arXiv, bioRxiv, SSRN, and PMC. Their conservative estimate is 146,932 hallucinated citations for 2025 alone, with an observed mid-2024 inflection; for the bioRxiv-to-PMC pairing they report 85.3% preprint-to-published persistence. The paper describes “real citations deployed to support claims the cited references do not actually make” as an open challenge. ARS v3.7.1 added trust-chain frontmatter for source provenance; v3.7.3 added locator infrastructure (three-layer citation anchors) for future claim-level audits and surfaces advisory risk signals at cite time (ARS labels the claim-faithfulness gap internally as “L3”; this is ARS terminology, not the paper’s). v3.7.x is motivated by Zhao et al.’s corpus-scale findings; corpus-scale evaluation of ARS itself remains future work.
v3.8 closes the second half of the L3 gap. v3.7.3 made every citation carry a locator anchor; v3.8 adds an opt-in audit pass (
ARS_CLAIM_AUDIT=1) that fetches the cited source against each anchor and judges whether the claim is actually supported. Five new HIGH-WARN classes (claim-not-supported, negative-constraint-violation, fabricated-reference, anchorless, constraint-violation-uncited) gate-refuse output through the formatter terminal hard gate. Calibration is shipped as a 20-tuple gold set with FNR<0.15 + FPR<0.10 acceptance thresholds; ramp-on plan is deferred to post-calibration evidence per v3.8 spec §5.Ren et al. (2026, Self-Improvements in Modern Agentic Systems: A Survey) supplies a third, survey-level anchor. Its scientific-discovery synthesis (§7.4) concludes that discovery agents cannot easily verify novelty, correctness, or reproducibility on their own and may exploit weak proxies instead, must manage evidence across heterogeneous tools and literature, and raise governance issues — “scientific writing can also amplify misinformation when the evidence is weak.” Its generation-loop chapters (§5.1–§5.2) list human auditing and retained human anchors among the practical safeguards for self-generated evaluation loops, and its historical chapter (§2.2) records the oldest form of the same lesson: the practical success of Lenat’s EURISKO depended heavily on the user serving as the external evaluation signal, pruning unproductive heuristic drift — a limitation the survey notes persists in modern agentic systems. ARS cites the survey as design rationale for its human-in-the-loop stance, not as empirical proof that human-in-the-loop pipelines outperform autonomous ones; the survey’s actionable deltas for ARS are tracked in #539–#541 and #547–#550.
v3.3 was inspired by PaperOrchestra (Song, Song, Pfister & Yoon, 2026, Google): Semantic Scholar API verification, anti-leakage protocol, VLM figure verification, and score trajectory tracking.
Architecture & pipeline
👉 docs/ARCHITECTURE.md — the full pipeline view: flow diagram, stage-by-stage matrix, data-access flow, skill dependency graph, quality gates, and mode list.
The architecture doc supersedes the sprawling pipeline description that used to live here. Everything about what runs in which stage now lives in one place.
Quick install
Prerequisites
ANTHROPIC_API_KEYexported, or set on firstclauderunPreToolUsewrite-scope guard (optional subagent hardening — if no real Python is found it cleanly no-ops and the guard is simply inactive; core skills are unaffected), plus a few opt-in features that shell out to Python (revision-patch mode, the submission-package verifier, and the/ars-cache-invalidate//ars-mark-read//ars-unmark-readcommands). On Windows, note thatpython3is often a non-functional Microsoft Store placeholder rather than real Python; install Python from python.org (or viawinget) so the launcher can find a real interpreter. The guard launcher is a POSIX shell script andhooks.jsoninvokes it throughbash, so on Windows it needs Git Bash (bundled with Git for Windows). With Git Bash present, a missing real Python degrades cleanly (the guard no-ops, silently). Without Git Bash, Claude Code falls back to PowerShell, which cannot run the.shlauncher at all: the guard is inactive and thePreToolUsehook will log an error per call rather than no-op quietly (accepted degradation — the guard is optional and never blocks your writes, but the hook noise is the trade-off until Git Bash is installed).Plugin install (v3.7.0+, recommended):
Verify it works: run
/ars-planand describe a paper you’re working on — ARS will start a Socratic dialogue to map out chapter structure. For a single-shot test instead, try/ars-lit-review "your topic".👉 docs/SETUP.md — full guide: install Claude Code, set up API keys, optional Pandoc/tectonic for DOCX/PDF, cross-model verification (
ARS_CROSS_MODEL), and six installation methods (Plugin, project skills, global skills, claude.ai Project, repo-cloned, Claude Science import).Using Claude Science? The four skills import directly: Skills → Import from GitHub, paste
https://github.com/Imbad0202/academic-research-skills, Preview, then Import 4 skills (requires v3.14.0+ of this repo — the importer reads the explicit skill paths in the marketplace manifest). Imports are point-in-time snapshots: re-import after ARS updates. Imported skills carry the ARS methodology (research / writing / review protocols); Claude Code-specific machinery — slash commands, hooks, subagent orchestration — does not transfer. See docs/SETUP.md Method 5 for details.Using Codex CLI? Install the sibling distribution instead:
Imbad0202/academic-research-skills-codex— same workflow content, Codex-native packaging as a single$academic-research-suiteskill withars-*aliases.Third-party platforms and integrations that wrap or host ARS are listed in THIRD_PARTY.md — community-submitted and not reviewed or endorsed by the maintainer.
Performance & cost
👉 docs/PERFORMANCE.md — per-mode token budgets, full-pipeline estimate (~$4–6 for a 15k-word paper), and recommended Claude Code settings (Auto mode; Agent Team optional).
Guides & articles
Features at a glance
repro_lock, optional cross-model integrity verification, mid-conversation reinforcement, and score trajectory tracking.data_access_level(raw/redacted/verified_only); enforced byscripts/check_data_access_level.py. Pattern adapted from Anthropic’s automated-w2s-researcher (2026). Seeshared/ground_truth_isolation_pattern.md.task_type(open-endedoroutcome-gradable). All current ARS skills areopen-ended.shared/benchmark_report_pattern.md.repro_locksub-block on Material Passport. Configuration documentation, not replay guarantee — LLM outputs are not byte-reproducible. Seeshared/artifact_reproducibility_pattern.md.ARS_MODEL_TIERINGswitch with two directions:economy(execution-type agents dispatch one tier below the session model, floor Opus-class) andquality-boost(judgment-type agents at integrity gates and final review step up to the frontier tier). Default unset = byte-equivalent to pre-#517 behavior. Seeshared/model_tiering.md.[CROSS-MODEL-HANDOFF v1]envelope with a normative Python grammar (scripts/cross_model_handoff.py) instead of prose-only enforcement, pinning agreement/divergence/malformed-result routing across all three checkpoint owners. Seeshared/cross_model_verification.md§”Cross-model handoff envelope”.experiment_provenance[]on the Material Passport records experiments the scholar ran externally (ARS never runs experiments), and manuscript claims join to them viaclaim_intent_manifest.planned_experiment_ids[]. The integrity gate (Stage 2.5/4.5) audits each experiment-backed claim against declared provenance —ALIGNED/OVERSTATED/NOT_SUPPORTED_BY_PROVENANCE/PROVENANCE_INSUFFICIENT— without judging whether the experiment itself was correct. A fail-closedexperiment_intake_declarationmakes “did you run experiments?” an explicit Stage 1 decision (even literature-only runs declareno_experiments_declared). Seeshared/handoff_schemas.md§”Experiment Provenance Intake (#260)”.Showcase: real pipeline output
See the complete artifacts from a real 10-stage pipeline run — peer review reports, integrity verification reports, and the final paper:
Browse all pipeline artifacts →
Companion: Experiment Agent
If your research involves running experiments (code or human studies) before writing, the Experiment Agent skill fills the gap between ARS Stage 1 (RESEARCH) and Stage 2 (WRITE).
What it does: executes code experiments (Python, R, etc.) with real-time monitoring, manages human study protocols with IRB ethics checklist, interprets statistics with 11-type fallacy detection, and verifies reproducibility.
How to use together: pause the ARS pipeline after Stage 1, run experiments in a separate experiment-agent session, then bring the results (with Material Passport) back to ARS Stage 2. ARS requires zero modification. See the experiment-agent README for setup instructions.
Stage 1 intake declaration (#260): at Stage 1, ARS detects whether the run will carry experiment-backed claims and sets a fail-closed
experiment_intake_declarationon the Material Passport. If you ran experiments externally, the scholar enters oneexperiment_provenance[]entry per experiment (experiment_id, nestedrepro_lock,planned_vs_executed[],negative_results[],known_limitations[]) and the declaration is set toexperiments_declared; if not, it is set tono_experiments_declared. The declaration is required on every post-#260 passport — a run that touches no experiments still declaresno_experiments_declared, so the integrity gate can never be silently bypassed by a forgotten provenance block. Theexperiment_ids are frozen at this intake point; the writers later reference them viaplanned_experiment_ids[].Teaching-side companion: Teaching Skills applies the ARS architecture (skill ensembles, shared contracts, staged gates, a Course Passport) to the teaching side of academic life — course design → lessons → assessment → delivery → reflection; its
sotlmode hands classroom-inquiry projects off to ARS deep-research / academic-paper for the publication phase.Usage
Quick Start
Individual Skills
Deep Research (8 modes)
Academic Paper (11 modes)
Academic Paper Reviewer (6 modes)
Academic Pipeline (Orchestrator)
Supported Languages
Supported Citation Formats
Supported Paper Structures
Skill Details
Per-agent responsibilities and per-stage artifacts now live in
docs/ARCHITECTURE.md. Version numbers are anchored here so release metadata stays in one place.Deep Research (v2.11.0)
13-agent research team. Modes: full, quick, review, lit-review, three-way-scan, fact-check, socratic, systematic-review. Full agent roster and artifacts: see ARCHITECTURE.md §3.
Academic Paper (v3.2.0)
12-agent paper writing pipeline. Modes: full, plan, outline-only, revision, revision-coach, abstract-only, lit-review, format-convert, citation-check, disclosure, rebuttal-audit. Output: MD + DOCX (via Pandoc when available) + LaTeX (APA 7.0
apa7class / IEEE / Chicago) → PDF via tectonic. Full agent roster and per-phase responsibilities: see ARCHITECTURE.md §3.Academic Paper Reviewer (v1.10.0)
7-agent multi-perspective review with 0-100 quality rubrics. Modes: full, re-review, quick, methodology-focus, guided, calibration. Decision mapping: ≥80 Accept, 65-79 Minor Revision, 50-64 Major Revision, <50 Reject. First-round review team vs. narrow re-review team boundary: see ARCHITECTURE.md §3 Stage 3 / Stage 3’.
Academic Pipeline (v3.19.0)
10-stage orchestrator with integrity verification, two-stage review, Socratic coaching, and collaboration evaluation. Pipeline guarantees: every stage requires user confirmation checkpoint; integrity verification (Stage 2.5 + 4.5) cannot be skipped; R&R Traceability Matrix (Schema 11) independently verifies author revision claims. v3.4 added the Compliance Agent (PRISMA-trAIce + RAISE) at Stage 2.5 / 4.5. v3.5 adds the Collaboration Depth Observer (
collaboration_depth_agent, advisory only — never blocks) at every FULL/SLIM checkpoint and at pipeline completion. MANDATORY integrity gates (2.5 / 4.5) explicitly skip the observer so compliance checks are not diluted. Based on Wang & Zhang (2026), IJETHE 23:11. Stage-by-stage matrix with agents, artifacts, and gates: see ARCHITECTURE.md §3.v3.0 Optimizations: What We Discovered About AI’s Structural Limits
What happened
While using ARS to write a reflection article about AI in higher education, I ran into three structural problems that no amount of prompt engineering could fix:
Frame-lock: I asked the AI to run a devil’s advocate debate against its own thesis. It did — four rounds, each more refined than the last. But every round stayed inside the frame I’d set. The DA attacked arguments, never premises. It never asked “are we even discussing the right question?” This is the same pattern that caused the 31% citation error rate in v2.7’s stress test: the verifying AI and the generating AI share the same cognitive frame.
Sycophancy under pushback: Every time I challenged the DA’s attacks, it conceded too quickly. It retracted findings faster than it launched them. The model’s training rewards conversational harmony — so “the user pushed back” was treated as evidence that the attack was wrong, when often it just meant the user was persistent.
Intent misdetection: The Socratic Mentor kept trying to converge and produce deliverables (“Want me to write this up?”) when I was still exploring. It couldn’t distinguish “the user wants a deep philosophical discussion” from “the user wants an RQ brief.” Both look like engagement, but they need opposite AI behaviors.
What we changed (v3.0)
Devil’s Advocate — Concession Threshold Protocol (
deep-research+academic-paper-reviewer)Socratic Mentor — Intent Detection Layer (
deep-research)Socratic Mentor — Dialogue Health Indicator (
deep-research)Why this matters
These optimizations don’t solve AI’s structural limits — they make the limits visible and manageable. The DA will still eventually concede if pushed hard enough. The Socratic Mentor will still have some convergence bias. But now there are explicit checkpoints that slow down the sycophancy, force the DA to justify concessions, and prevent the Mentor from wrapping up before the user is ready.
The deeper lesson: AI literacy isn’t about learning to use AI as a tool, following ethics rules, or fearing AI risks. It’s about engaging AI deeply enough to discover its structural limits yourself — and your own thinking limits in the process.
License
This work is licensed under CC-BY-NC 4.0.
You are free to:
Under the following terms:
Attribution format:
Contributors
Cheng-I Wu (吳政宜) — Author and maintainer
aspi6246 — Contributor. The v3.1 optimization was inspired by patterns from Claude-Code-Skills-for-Academics: read-only constraint pattern, anti-pattern codification as first-class design, cognitive framework approach (teaching “how to think” not just procedures), and lean skill size philosophy.
mchesbro1 — Contributor. Originally proposed and drafted the IS Basket of 8 journals for
academic-paper-reviewer/references/top_journals_by_field.md(Issue #5).cloudenochcsis — Contributor. Extended the IS section from the Basket of 8 to the full Senior Scholars’ Basket of 11 — adding Decision Support Systems, Information & Management, and Information and Organization (Issue #7, PR #8). Sourced from the AIS Senior Scholars’ List of Premier Journals.
eltociear (Ikko Eltociear Ashimine) — Contributor. Translated the Japanese README (
README.ja-JP.md) (PR #161).xpfo-go (xpfo) — Contributor. Translated the Simplified Chinese README (
README.zh-CN.md) (PR #181).devCharlotte — Contributor. Translated the Korean README (
README.ko-KR.md) (PR #469).Yaobin29 — Contributor. Proposed reviewer-response tooling in PR #433; the
deep-research three-way-scanmode and theacademic-paper rebuttal-auditmode (rescued from the PR’sauditconcept) were integrated from that contribution in v3.12.1.Changelog
v3.19.0 (2026-07-22) — Revision-round claim-drift guards, PDF read-integrity preflight, read-scope attestation
v3.18.0 (2026-07-18) — Self-improvement survey integration: advisory quality layers, risk-stratified claim gate, cross-model reviewer & judge tracks
v3.17.0 (2026-07-16) — Pipeline boundary semantics, canonical cross-model handoff envelope, executable panel checker
v3.16.0 (2026-07-12) — Model tiering, cross-model gate hardening, WP advisory sharpening
v3.15.0 (2026-07-04) — Release-gate hardening, prompt-debt retirement round 2, defrift locks
v3.14.0 (2026-07-02) — Claude Science importability, eval-comment rendering, prompt-debt retirement
v3.13.0 (2026-06-18) — Hook portability, provider-agnostic verification, guard correctness
v3.12.1 (2026-06-15) — Reviewer-response triage modes (PR #433 integration)
v3.12.0 (2026-06-08) — Kong auto-research feature track: experiment provenance, figure fidelity, cross-paper contradiction, partial-evidence decomposition
v3.11.1 (2026-06-06) — Post-ship correctness, hardening & provenance rollup
v3.11.0 (2026-06-04) — Deterministic citation verification gate (#182)
v3.10.0 (2026-06-01) — Triangulation policy layer, Kong survey adoptions, eval harness, scoped-write guard
v3.9.4.2 (2026-05-19) — post-ship hotfix for PR #149 CI discipline gates (codex post-ship)
v3.9.4.1 (2026-05-19) — post-ship hotfix for v3.9.4 temporal verification (#135 codex post-ship)
v3.9.4 (2026-05-18) — #135 temporal verification layer (advisory)
v3.9.3 (2026-05-18) — #128 housekeeping (shared client utilities + dedup resolvers)
v3.9.2 (2026-05-18) — #133 phase boundary hot-fix
v3.9.1 (2026-05-18) — #129 + #130 client hardening
v3.9.0 (2026-05-17) — #102 cross-index triangulation measurement
Migration: v3.7.3 corpora — run
python scripts/migrate_literature_corpus_to_v3_9_0.py PATHto backfill the two new fields. Pre-v3.7.3 corpora — runmigrate_literature_corpus_to_v3_7_3.pyFIRST, then v3.9.0 migration (daisy-chained per spec §3.7; the v3.9.0 tool only acts on entries that already carrycontamination_signals.semantic_scholar_unmatched).v3.8.2 (2026-05-17) — #118 uncited audit_tool_failure surface
uncited_audit_failure.schema.jsonaggregate (spec §3.6). One entry per uncited sentence × manifest pair where the constraint judge raisedJudgeInvocationError. Same fault-class enum as cited-path INV-14 (judge_timeout/judge_api_error/judge_parse_error/cache_corruption/retrieval_api_error/retrieval_timeout/retrieval_network_error).rule_version: D4-c-v1-uaf-v1.finding_iduniqueness, scoped_manifest_id cross-array integrity, (M, C) pair integrity when manifest_claim_id non-null, per-(sentence, manifest) dedup, rationale fault_class prefix, cross-aggregate exclusivity vsconstraint_violations[].[CLAIM-AUDIT-TOOL-FAILURE-UNCITED — <fault-class>], gate passes (retry-next-pass remediation). Formatter REFUSE list unchanged — UAF is advisory.scripts/claim_audit_pipeline.py): swallow site at line 1211-1224 removed;JudgeInvocationErrornow emits a UAF row +continues to the next (sentence, manifest) pair. No fake NOT_VIOLATED reachesconstraint_violations[].academic-pipeline/agents/claim_ref_alignment_audit_agent.md): Output emission table grows seventh row; Error handling table grows from 3 surfaces to 4 surfaces with the uncited-path UAF row.v3.8.0 (2026-05-16) — L3 Claim-Faithfulness Locator + Audit (paired milestone)
claim_ref_alignment_audit_agent(v3.8 PR #121). Opt-in (ARS_CLAIM_AUDIT=1, default OFF) Stage 4→5 audit agent. Judges every sampled citation against retrieved excerpt; emitsclaim_audit_results[]+claim_intent_manifests[]+claim_drifts[]+uncited_assertions[]+constraint_violations[]aggregates. 8-row finalizer matrix routes HIGH-WARN classes (CLAIM-NOT-SUPPORTED / NEGATIVE-CONSTRAINT-VIOLATION / FABRICATED-REFERENCE / ANCHORLESS / CONSTRAINT-VIOLATION-UNCITED) through the formatter REFUSE rules 6-10. Calibration runner ships with 20-tuple gold set (T-C1 FNR<0.15 + FPR<0.10, T-C2 per-class, T-C3 shape integrity). 8 rounds of dual-track review (R1 codex + Gemini-3.1-pro-preview, R2-R8 codex-only after Gemini quota exhausted); trajectory R1 4P1+2P2 → R8 0P1+4P2 ship gate.synthesis_agent/draft_writer_agent/report_compiler_agentgain## Three-Layer Citation Emission (v3.7.3)H2. Every<!--ref:slug-->carries<!--anchor:<kind>:<value>-->with<kind> ∈ {quote, page, section, paragraph, none}(quote anchors capped at 25 words, URL-encoded).pipeline_orchestrator_agentfinalizer becomes 5-cell with precedence-zero NO-LOCATOR check.formatter_agentadds explicit hard-gate refusal for[UNVERIFIED CITATION — NO QUOTE OR PAGE LOCATOR].literature_corpus_entry.schema.jsonadds optionalcontamination_signals: { preprint_post_llm_inflection, semantic_scholar_unmatched }object.bibliography_agentcomputes both signals at ingest. 11-round review trajectory (Codex×10 + Gemini cross-model×1) closed 22 findings. Spec:docs/design/2026-05-12-ars-v3.7.3-claim-faithfulness-and-contaminated-source-spec.md. External motivation: Zhao et al. arXiv:2605.07723 (2026-05).slr_lineageemission on systematic-review → academic-paper handoff (2026-05-15). Schema 9 optional booleanslr_lineagefield; producerpipeline_orchestrator_agentwrites at every handoff transition; consumerdisclosuremode dispatches--policy-anchor=prisma-trAIceper the §4.3 G2 invariant track gate.README.zh-TW.mdmotivation section frames the v3.7.x line against Zhao et al.’s 146,932 hallucinated-citation finding.scripts/migrate_literature_corpus_to_v3_7_3.pyretro-computes both contamination signals across pre-v3.7.3 passports.scripts/semantic_scholar_client.pyadds 1-req/s throttle (drops to 0.1s whenS2_API_KEYdetected), outage latch on URLError, andreset_outage_latch()for long-running cross-passport batches.v3.7.0 (2026-05-05) — Claude Code Plugin Packaging
.claude-plugin/plugin.jsondeclares the suite (4 skills auto-discovered fromskills/directory via relative symlinks)..claude-plugin/marketplace.jsonregisters the plugin so a single GitHub-hosted endpoint serves both the marketplace listing and the plugin source. README +README.zh-TW.md+docs/SETUP.mdcarry dual-track install instructions.commands/ars-*.md(Phase 2.1, PR #69) mappingMODE_REGISTRY.mdentries to/ars-<mode>triggers. Model routing is pinned in each command’s frontmatter —opusforfullandrevision-coach(architectural / review-interpretation depth),sonnetfor the other 8. No Haiku per project policy.agents/*_agent.md(Phase 2.1, PR #69) as relative symlinks to the v3.6.7-hardened downstream agents indeep-research/agents/:synthesis_agent,research_architect_agent,report_compiler_agent. Underscore filenames preserved to keepscripts/check_v3_6_7_pattern_protection.pyhard-pinned paths and INV-3 manifest-confined Clause 1 invariant intact. Symlinks (not copies) preserve a single source of truth and prevent the Pattern C3 attack surface that v3.6.7 §6 inversion sweep + INV-1/2/3 lint closes. (Materialized to real byte-identical copies in #413 — relative symlinks break Windows checkouts withoutcore.symlinksand zip-download installs; the single-source guarantee moved to thescripts/check_agents_mirror_sync.pybyte-equality CI lint.)model: inheritadded to those three source agent frontmatters. Inherit chosen over pinningsonnetso an opus session running ARS full pipeline keeps opus agents (instead of being capped). The user’s~/.claude/hooks/warn-agent-no-model.shPreToolUse hook gates Haiku at the dispatching boundary, soinheritresolves through an already-Haiku-free model.hooks/hooks.json+scripts/announce-ars-loaded.sh(Phase 2.2, PR #70). When the plugin loads, the hook injects anadditionalContextlisting the 10 slash commands, the 3 plugin agents, and a token-budget pointer into the LLM’s first turn.startupandclearsource values get the full announce;resumeandcompactget a one-line ack to avoid burning context. Bash 3.2 compatible — runs on macOS stock/bin/bashwith nobrew install bashrequirement.SubagentStop → run_codex_audit.shcodex audit hook was scoped out for v3.7.0 due to a contract gap (the SubagentStop payload carries no stage/deliverable info, so the wrapper would have to half-infer required arguments) and an invoker-class boundary (run_codex_audit.shlines 4–7 forbid same-session in-LLM invocation; PostToolUse fires inside the producing session). Real audit-hook integration deferred to a future release when ARS gains a stage/deliverable propagation contract. Seedocs/design/2026-04-30-ars-v3.7.0-plugin-packaging-roadmap.mdUpdate note 2026-05-05 (Phase 2.2 scope reduction).docs/PERFORMANCE.md+.zh-TW.mdgain a “v3.7.0 Plugin agents and model routing” subsection explaining the inherit semantics and current 3-agent scope boundary.${CLAUDE_PLUGIN_ROOT}breaking install paths with spaces) that the inline rounds missed — confirms the value of separating implementation review (inline) from contract review (fresh).commands/,agents/,hooks/,.claude-plugin/,skills/symlink dir, three plugin-agentmodel: inheritfrontmatter additions). Existing 4.3k clone-install users see no breaking change.v3.6.8 (2026-05-03) — Generator-Evaluator Contract Gate (v3.6.6 spec ship)
shared/sprint_contract.schema.json) extends Schema 13 with two newmodeenum values (writer_full+evaluator_full), two new optional top-level fields (pre_commitment_artifactswriter-only,disagreement_handlingevaluator-only), and 12allOfbranches enforcing reviewer- / writer- / evaluator-conditional gates. Existing reviewer contracts validate byte-equivalent under Schema 13.1 (§3.6 zero-touch promise).shared/contracts/writer/full.json(D1–D7, F1/F4/F2/F3/F0) andshared/contracts/evaluator/full.json(D1–D5, F1/F2/F3/F6/F4/F5/F0). Promoted from design-time artefacts on the spec branch to live shipped status atomically with the Schema 13.1 upgrade.academic-paper full: Phase 4 splits into Phase 4a (writer paper-blind pre-commitment) + Phase 4b (writer paper-visible drafting + self-scoring); Phase 6 splits into Phase 6a (evaluator paper-blind pre-commitment) + Phase 6b (evaluator paper-visible scoring + decision). Phase-numbered<phase4a_output>/<phase6a_output>data delimiters mirror the v3.6.2 reviewer pattern. Lint count summary: writer 3+4 / evaluator 5+5 / reviewer 5+6 (reviewer remains zero-touch).academic-paperSKILL + agent files gain a verbatim## v3.6.6 Generator-Evaluator Contract Protocolblock (101 lines in SKILL.md plus 47 lines indraft_writer_agent.md+ 57 lines inpeer_reviewer_agent.md). SKILL.md also adds a new## Known limitationssection carrying graceful-degradation + cross-session resume forward notes for v3.6.7+.scripts/check_sprint_contract.pySC-* mode-gating audit (SC-5 + SC-11 reviewer-only; SC-9 extended across all three mode families). 17 new tests bring the validator unit-test count from 54 to 71 (positive + 5 schema-branch negative + 2 §3.6 reviewer regression + 6 mode-gating tests).scripts/check_v3_6_6_ab_manifest.pyenforces §6.2 manifest schema + §6.5 git-tracked invariants ontests/fixtures/v3.6.6-ab/manifest.yaml..github/workflows/spec-consistency.ymlextends the sprint contract validation loop to iterate writer + evaluator template directories alongside the existing reviewer loop, plus runs the new manifest CI lint.tests/fixtures/v3.6.6-ab/(30 files): manifest + README + 6 paper-A inputs/baseline + 1 paper-C inputs/baseline + Stage 3 reviewer excerpt + 6 codex-judge baseline placeholders. Real fixture data populates in follow-up commits before the implementation work fully completes.v3.6.7 (2026-04-30) — Downstream-Agent Pattern Protection (Step 1+2)
synthesis_agent(A1–A5 narrative-side), the survey-designer mode ofresearch_architect_agent(B1–B5 instrument-side), and the abstract-only mode ofreport_compiler_agent(C1–C3 publication-side). Each agent prompt now carries aPATTERN PROTECTION (v3.6.7)block.shared/references/:irb_terminology_glossary.md,psychometric_terminology_glossary.md,protected_hedging_phrases.md,word_count_conventions.md. The reference files carry operational contracts that the agent prompts cite by path.shared/templates/codex_audit_multifile_template.mdwith seven audit dimensions and a mandatory three-part Section 4(f) check forreport_compiler_agentbundles. Failure of any sub-check is a P1 finding.scripts/check_v3_6_7_pattern_protection.pyenforces protection-clause presence and obligation-phrase shape;scripts/test_check_v3_6_7_pattern_protection.pypreserves codex review evidence so future checker regressions surface in CI. Both are wired into.github/workflows/spec-consistency.yml.gpt-5.5+xhighcross-model review reached SHIP-OK with zero P1+P2 findings. Step 6 (orchestrator runtime hooks) and Step 8 (synthetic eval case) ship in a follow-up PR.v3.6.5 (2026-04-27) — Material Passport
literature_corpus[]Consumer Integrationdeep-research/agents/bibliography_agent.mdandacademic-paper/agents/literature_strategist_agent.md. Both follow the same five-step corpus-first, search-fills-gap flow when the passport carries a non-emptyliterature_corpus[]and the same four Iron Rules (Same criteria / No silent skip / No corpus mutation / Graceful fallback on parse failure).obtained_via/obtained_at.final_included = pre_screened_included[] ∪ external_included[]stays neutral — no provenance tags on bibliography entries or literature matrix rows.academic-pipeline/references/literature_corpus_consumers.mdwith the canonical PRE-SCREENED template, BAD/GOOD examples, four Iron Rules, and per-consumer reading instructions.scripts/check_corpus_consumer_protocol.pyenforcing nine protocol invariants with manifest-driven consumer list (scripts/corpus_consumer_manifest.json).shared/handoff_schemas.mdretired the v3.6.4 “Consumer-side integration deferred to v3.6.5+” caveat; replaced with backpointer to the consumer protocol.[CORPUS PARSE FAILURE]surface.citation_compliance_agentcorpus integration deferred (target version TBD post-v3.8).v3.6.4 (2026-04-25) — Material Passport
literature_corpus[]Input Portliterature_corpus[]field added to Schema 9 as an optional input port for user-owned literature. Each entry conforms toshared/contracts/passport/literature_corpus_entry.schema.json(CSL-JSON authors, year, title, source_pointer + private optionalabstract/user_notes).academic-pipeline/references/adapters/overview.md: any program (any language) reading a user corpus source can produce conformantpassport.yaml+rejection_log.yaml. Fail-soft entry-level errors, fail-loud adapter-level errors, deterministic ordering.scripts/adapters/:folder_scan.py(filesystem of PDFs),zotero.py(Better BibTeX JSON export),obsidian.py(vault frontmatter). Starting points only; users are expected to write their own adapters for non-reference sources.shared/contracts/passport/rejection_log.schema.jsonwith closed enum of categorical reason values; always emitted (empty when no rejections).scripts/check_literature_corpus_schema.pyvalidates schemas + adapter examples;scripts/sync_adapter_docs.py --checkprevents schema→docs drift; newpytest.ymlworkflow runsscripts/adapters/tests/on path-filtered triggers.bibliography_agentandliterature_strategist_agentwere wired in v3.6.5.v3.6.3 (2026-04-23) — Opt-in Passport Reset Boundary
ARS_PASSPORT_RESET=1). Promotes every FULL checkpoint to a context-reset boundary. Newresume_from_passport=<hash>mode lets users resume in a fresh Claude Code session from the Material Passport ledger alone.systematic-reviewmode with the flag ON makes reset mandatory at every FULL checkpoint; other modes treat reset as the flag-gated default. Flag OFF preserves pre-v3.6.3 behavior byte-for-byte.reset_boundary[]ledger with two entry kinds (kind: boundary+kind: resume). Hash uses JSON Canonical Form + SHA-256 with canonical placeholder for self-reference safety. Optionalpending_decisionhandles MANDATORY branch choices.scripts/check_passport_reset_contract.pyCI lint: every mention of the flag must co-locate a pointer to the authoritative protocol doc.academic-pipeline/references/passport_as_reset_boundary.md.docs/PERFORMANCE.mdupdated with long-running-session guidance.v3.6.2 (2026-04-23) — Reviewer Sprint Contract Hard Gate
v3.6.2 introduces Schema 13 sprint contracts and a hard-gate orchestration that forces reviewers to pre-commit their scoring plan before reading the paper. Reviewer-only first test case; writer/evaluator deferred to v3.6.4. See CHANGELOG.
panel_size,acceptance_dimensions,failure_conditions(withseverityprecedence + panel-relativecross_reviewer_quantifier),measurement_procedure, optionaloverride_ladder, boundedagent_amendments. Validator:scripts/check_sprint_contract.py.<phase1_output>...</phase1_output>data delimiter to narrow the self-injection surface.failure_conditionwith panel-relative quantifier + recognised expression vocabulary → resolve precedence byseverity. Forbidden-ops list explicit ineditorial_synthesizer_agent.shared/contracts/reviewer/full.jsonpanel 5;shared/contracts/reviewer/methodology_focus.jsonpanel 2).reviewer_re_review,reviewer_calibration,reviewer_guidedare reserved in the schema enum but ship without contract templates in v3.6.2; they retain pre-v3.6.2 behaviour.reviewer_quickis excluded from the enum entirely.academic-paper-reviewerSKILL version:1.8.1 → 1.9.0.academic-pipelineSKILL version:3.5.1 → 3.6.2(suite-version invariant). Suite version bumped to3.6.2.docs/design/2026-04-23-ars-v3.6.2-sprint-contract-design.mdand protocolacademic-paper-reviewer/references/sprint_contract_protocol.md.v3.5.1 (2026-04-22) — Opt-in Socratic Reading-Check Probe
v3.5.1 adds an opt-in honesty probe to the Socratic Mentor (
ARS_SOCRATIC_READING_PROBE=1). Default off. See CHANGELOG.ARS_SOCRATIC_READING_PROBE=1is set, the Socratic Mentor fires a one-time honesty probe during goal-oriented sessions where the user has cited a specific paper. Decline is logged without penalty. Outcome flows into the Research Plan Summary and Stage 6 AI Self-Reflection Report. No new agent, no schema change.deep-researchSKILL version:2.9.0 → 2.9.1.academic-pipelineSKILL version:3.5.0 → 3.5.1. Suite version bumped to3.5.1.v3.5.0 (2026-04-21) — Collaboration Depth Observer
collaboration_depth_agentinacademic-pipeline(Agent Team grows from 3 to 4). Invoked at every FULL/SLIM checkpoint and at pipeline completion; scores user-AI collaboration against a 4-dimension rubric. Advisory only — never blocks progression. MANDATORY checkpoints (Stages 2.5 / 4.5 integrity gates) do NOT invoke the observer.shared/collaboration_depth_rubric.mdv1.0. Dimensions: Delegation Intensity, Cognitive Vigilance, Cognitive Reallocation, Zone Classification (Zone 1 / Zone 2 / Zone 3). Based on Wang, S., & Zhang, H. (2026). “Pedagogical partnerships with generative AI in higher education: how dual cognitive pathways paradoxically enable transformative learning.” International Journal of Educational Technology in Higher Education, 23:11. DOI 10.1186/s41239-026-00585-x.ARS_CROSS_MODELis set the observer runs on both models; dimension disagreement > 2 points is reported rather than silently smoothed.ARS_CROSS_MODEL_SAMPLE_INTERVALescape hatch for cost trade-off.insufficient_evidenceblock instead of dispatching the full-model observer.academic-pipelineSKILL version:3.3.0 → 3.4.0. Suite version bumped to3.5.0. New lintscripts/check_collaboration_depth_rubric.py+ 10 tests.v3.4.0 (2026-04-20) — Compliance Agent + Schema 12
compliance_history[](append-only).disclosure_addenduminto manuscript. No detection evasion possible.task_type: open-ended.v3.3.6 (2026-04-15) — README Streamlining + ARCHITECTURE doc
docs/ARCHITECTURE.mdas the single source of truth for pipeline structure (flow, matrix, data-access, dependency graph, quality gates, modes). Merged into main via PR #18.docs/SETUP.md(prerequisites, API keys, Pandoc/tectonic, cross-model verification, installation methods) anddocs/PERFORMANCE.md(token budgets, recommended Claude Code settings). README links to both instead of inlining them.3.3.6.v3.3.5 (2026-04-15)
benchmark_report.schema.json+repro_lockoptional block on Material Passport. Both ship with pattern docs, lints, and examples. First formal Python dev dep manifest (requirements-dev.txt).v3.3.4 (2026-04-15) — README Changelog Sync Patch
Synced the embedded changelog sections in
README.mdandREADME.zh-TW.mdso they include the missingv3.3.3andv3.3.2release summaries.Extended
scripts/check_spec_consistency.pyso future README changelog drift fails CI.v3.3.3 (2026-04-15) — Release Prep + Lint Hardening
Hardened SKILL frontmatter linting: missing closing
---fences now fail cleanly instead of being parsed as valid YAML.Frontmatter that parses as valid YAML but not as a mapping now reports a readable error instead of crashing.
Fixed the broken showcase link for the post-publication audit report in both READMEs.
Added README relative-link validation to the spec consistency check so dead links fail CI.
Aligned the DOCX output contract across the docs: direct
.docxgeneration is Pandoc-dependent, with Markdown + conversion instructions as fallback.Prepared the
v3.3.3release: suite version bump,academic-paper-> v3.0.2,academic-pipeline-> v3.2.2.v3.3.2 (2026-04-15) — Data Access Levels + Task Type Metadata
metadata.data_access_levelto all top-levelSKILL.mdfiles with enforced vocabulary:raw,redacted,verified_only.metadata.task_typeto all top-levelSKILL.mdfiles with enforced vocabulary:open-ended,outcome-gradable.shared/ground_truth_isolation_pattern.mdand linked the new vocabulary fromshared/handoff_schemas.md.v3.3.1 (2026-04-14) — Spec Consistency Patch
.claude/CLAUDE.md,MODE_REGISTRY.md, andSKILL.mdfiles to the current mode counts and published skill versions.v3.3 (2026-04-09) — PaperOrchestra-Inspired Enhancements
Integrates techniques from PaperOrchestra (Song, Song, Pfister & Yoon, 2026, Google).
[MATERIAL GAP]for missing content instead of filling from memory. Reduces Mode 5/6 failure risk.v3.2 (2026-04-09) — Lu 2026 Nature Integration
Integrates insights from Lu et al. (2026, Nature 651:914-919) — the first end-to-end autonomous AI research system to pass blind peer review.
v3.1.1 (2026-04-09) — IS Senior Scholars’ Basket of 11
External contributions: @mchesbro1 originally proposed and drafted the IS Basket of 8 journals (Issue #5); @cloudenochcsis extended it to the full Senior Scholars’ Basket of 11 (Issue #7, PR #8). Updated
academic-paper-reviewer/references/top_journals_by_field.mdSection 7, adding Decision Support Systems, Information & Management, and Information and Organization. Source: AIS Senior Scholars’ List of Premier Journals.v3.1 (2026-04-06) — Anti-Context-Rot + Cognitive Frameworks + Lean Size
Inspired by patterns from aspi6246/Claude-Code-Skills-for-Academics.
Wave 1: Anti-Context-Rot Anchors
Wave 2: Traceability + Cognitive Frameworks + Reinforcement
argumentation_reasoning_framework.md— Toulmin model, Bradford Hill causal reasoning, inference to best explanation, epistemic status classificationreview_quality_thinking.md— three lenses (internal validity, external validity, contribution), common reviewer traps, calibration questionswriting_judgment_framework.md— clarity test, reader’s journey, discipline-specific voice, revision decision matrixWave 3: Lean Skill Size
references/filesv3.0 (2026-04-03) — Anti-Sycophancy + Intent Detection + Dialogue Health
ARS_CROSS_MODELenv var — without it, everything works as before. Seeshared/cross_model_verification.mdfor full setup guide, API patterns, and cost estimates.v2.9.1 (2026-04-03) — Skill Metadata
status: activeandrelated_skillscross-references to all 4 SKILL.md frontmatters.deep-research↔academic-paper↔academic-paper-reviewer↔academic-pipeline.v2.9 (2026-03-27) — Style Calibration + Writing Quality Check
shared/style_calibration_protocol.mdacademic-paper/references/writing_quality_check.md): Writing quality checklist applied during draft self-review. 5 categories: AI high-frequency term warnings (25 terms), punctuation pattern control (em dash ≤3), throat-clearing opener detection, structural pattern warnings (Rule of Three, uniform paragraphs, synonym cycling), and burstiness checks (sentence length variation). These are good writing rules — not detection evasionshared/handoff_schemas.md)v2.8 (2026-03-22) — SCR Loop Phase 1: State-Challenge-Reflect
deep-research/references/socratic_questioning_framework.md: SCR Overlay Protocol mapping SCR phases to Socratic functionsCHANGELOG.mdv2.7 (2026-03-09) — Integrity Verification v2.0: Anti-Hallucination Overhaul
v2.6.2 (2026-03-09) — Intent-Based Mode Activation
socratic/planoverfull— safer to guide first.v2.6.1 (2026-03-09) — Bilingual Trigger Keywords
v2.6 / v2.4 / v1.4 (2026-03-08) — 15+ Improvements
apa7document class, text justification fix (ragged2e+etoolbox), table column width formula, bilingual abstract centering, standardized font stack (Times New Roman + Source Han Serif TC VF + Courier New), PDF via tectonic onlyv2.4 / v1.3 (2026-03-08)
v2.3 / v1.3 (2026-03-08)
tectonic(no HTML-to-PDF); APA 7.0 usesapa7document class (manmode) with XeCJK for bilingual CJK support; font stack: Times New Roman + Source Han Serif TC VF + Courier Newv2.2 / v1.3 (2025-03-05)
v2.0.1 (2026-03)
v2.0 (2026-02)
integrity_verification_agent— 100% reference/data verification with audit traildevils_advocate_reviewer_agent— 8-dimension thesis challengerv1.0 (2026-02)