Current
Document-level self-contradiction detection
Benchmark first
Corpora: FEVER · FEVEROUS · ContraDoc · FreshQA
Protocol: Report verdict macro-F1 and accuracy separately from evidence precision/recall/F1 and contradiction localization. Pin corpus and source revisions. Replay corrections and removals, then measure stale-answer rate, time-to-consistency, unsupported-answer rate, and authorization leakage. A correct label without the complete evidence set does not pass the evidence contract.
This is not corpus retrieval. It validates whether one multi-sentence document is judged to contradict itself, where the conflict occurs, how much of the document the reasoning inspected, and how an external reinforcement-learning trainer should score the result.
How it works
Tag sentences. Number the document from 1 through
nbefore inference.Propose a judgment. An injected model returns a Boolean judgment, localized evidence sentence IDs, and reasoning containing
[i],[i-j], or[i]-[j]references.Validate localization. Mari expands ranges, rejects out-of-document references, requires evidence for positive judgments, and forbids contradiction evidence on negative judgments.
Measure reference coverage. Deduplicate every sentence mentioned in reasoning and compute
|S_covered| / |S_total|.Compute independent rewards. Return accuracy, reference-coverage, and format components for an external GRPO trainer. A correct positive judgment without any gold-evidence hit receives
-1; a correct localized judgment receives1 + matched/gold.
Tagged document[1] … [2] … [n]
model
Judgment + evidencereasoning references
validate
Assessmentlocalized + coverage
train/evaluate
Reward componentsaccuracy · coverage · format
from mari_components.verification import (
document_contradiction_rewards, validate_document_contradiction,
)
assessment = validate_document_contradiction(
sentence_count=len(sentences), judgment=proposal.judgment,
evidence_sentence_ids=proposal.evidence_sentence_ids,
reasoning=proposal.reasoning,
)
rewards = document_contradiction_rewards(
assessment, expected_judgment=case.is_self_contradictory,
gold_evidence_sentence_ids=case.conflicting_sentence_ids,
format_valid=proposal.matches_required_format,
)
What Mari does not claimReference coverage measures which sentence tags appeared in reasoning; it does not prove the reasoning is valid. Mari validates and scores a proposed judgment but does not replace the teacher-distilled SFT model, GRPO trainer, or semantic contradiction verifier.
Papers
Reinforced Reference Coverage for Document-Level Self-Contradiction DetectionContraDoc benchmark
Mari implements sentence-reference parsing, localization invariants, Equation 7 coverage, and Equations 5–8 reward components. These were checked against the MIT RRC-DSCD implementation and Apache-2.0 ContraDoc boundary. The RRC repository’s current accuracy code diverges from published Equation 5: it normalizes by predicted evidence and produces a 0.5 zero-hit score. Mari deliberately retains the paper’s gold-normalized term and -1 zero-hit result.