Score every claim in your LLM output against your knowledge base, with calibrated confidence and auditable evidence. NLI + RAG fact-checking, explicit abstention, prompt-injection detection, sealed evidence packets, and bounded streaming contradiction checks across a 530-module capability surface.
LLMs hallucinate. Your users trust them anyway. One wrong medical dosage. One fabricated legal citation. One invented financial figure. By the time a human reviewer catches it, the damage is done. Generic output filters catch obvious toxicity but miss subtle factual errors — the kind that sound perfectly plausible. Director-AI scores every claim against your knowledge base, returns an auditable verdict before the output is trusted, and — opt-in — halts streamed claims that contradict your grounding.
How it works
LLM Output
→
Claim Extraction
→
NLI Scoring FactCG 0.4B
→
RAG Fact-Check Your knowledge base
→
Dual Entropy Confidence + divergence
→
■ Halt stream
/
✓ Pass
Core features
Opt-in streaming contradiction check
Halts streamed claims that contradict your retrieved grounding facts. Response-level scoring is the production gate; the streaming check is opt-in and evidence-bound — not a sole guarantee.
Dual-entropy scoring
NLI contradiction detection (FactCG-DeBERTa, 0.4B params) combined with RAG fact-checking against your knowledge base. Two independent signals, one confidence score.
# Installpip install director-ai[all]
# Score a claim against a sourcefrom director_ai import score
result = score("The Earth is 4.5 billion years old", "The Earth formed approximately 4.54 billion years ago.")
print(result) # GuardResult(score=0.94, passed=True)# Or run as a REST proxy (zero code changes to your app)director-ai serve --port 8000 --upstream https://api.openai.com/v1
NLI models
FactCG-DeBERTa-v3-Large
Default scorer. 0.4B params, MIT licensed. Best speed/accuracy trade-off. ONNX + TensorRT GPU acceleration paths available.
MiniCheck-Flan-T5-L
0.8B params. Higher accuracy (77.4%) at ~3× latency cost. Best for offline batch verification.
MiniCheck-DeBERTa-L
0.4B params. Alternative DeBERTa backbone with different NLI training data.
Gemma 4 E4B (LLM-as-judge)
LLM-based scoring for complex claims. Highest accuracy but sends data to external provider. Off by default.
Heuristic-only (Lite)
Zero-dependency scorer using word overlap, numeric consistency, and structural checks. <0.5 ms. ~55% accuracy. CPU-only fallback.
Rust backend (backfire)
Native compiled compute via backfire-kernel. 12 accelerated functions. No Python GIL. No CUDA dependency for basic scoring.