Home▸LLM Hallucination Detection & Safety▸LLM Confabulation Risk by Query Complexity Simulator

⚠️ LLM Confabulation Risk by Query Complexity Simulator

This simulation examines how the complexity and length of queries influence the risk of an artificial intelligence generating inaccurate or fabricated information, providing insights into potential risks and mitigation strategies.

LLM Hallucination Detection & Safety2DModerate60 FPS
llm-confabulation-query-complexity-simulator ↗ Open standalone

Simple Queries Answered from Dense Training Coverage

Placeholder: common facts are well represented in training data.

  • 1 / 5: Query complexity (single-step lookup)
  • ~96%: Training coverage (placeholder estimate)
  • Low: Confabulation risk (placeholder estimate)
  • Minimal: Verification needed (placeholder guidance)

Why simple queries are low risk

Placeholder: dense coverage means the model has seen this pattern many times.

What reliable grounding looks like

Placeholder: high-frequency facts sit well inside the model's reliable-knowledge region.

Moderate Complexity Introduces the First Knowledge Gaps

Placeholder: bridging between known facts starts to strain grounding.

  • 2 / 5: Query complexity (placeholder)
  • ~80%: Training coverage (placeholder estimate)
  • Low–moderate: Confabulation risk (placeholder estimate)
  • Light spot-check: Verification needed (placeholder guidance)

Bridging between known facts

Placeholder: the model interpolates between two well-known facts.

Early signs of strain

Placeholder: fluent phrasing can mask small factual gaps here.

Complex Multi-Hop Queries Chain Several Reasoning Steps

Placeholder: each extra hop adds a chance of an unsupported leap.

  • 4 / 5: Query complexity (placeholder)
  • ~65%: Training coverage (placeholder estimate)
  • Moderate: Confabulation risk (placeholder estimate)
  • Step-by-step review: Verification needed (placeholder guidance)

Chained reasoning steps

Placeholder: each hop compounds uncertainty from the previous one.

Where the chain breaks

Placeholder: the weakest link sets the reliability of the whole chain.

Rare and Obscure Topics Force Interpolation

Placeholder: sparse training data sharply raises confabulation risk.

  • 4 / 5: Query complexity (placeholder)
  • ~22%: Training coverage (placeholder estimate)
  • High: Confabulation risk (placeholder estimate)
  • Full source check: Verification needed (placeholder guidance)

Interpolation past the boundary

Placeholder: the model extrapolates beyond reliable coverage.

Fluency without grounding

Placeholder: confident tone does not imply factual grounding here.

Confabulation Rate Rises Non-Linearly with Complexity and Rarity

Placeholder: hardest, highest-stakes queries carry the most risk.

  • 5 / 5: Query complexity (placeholder)
  • ~8%: Training coverage (placeholder estimate)
  • Critical: Confabulation risk (placeholder estimate)
  • Mandatory expert review: Verification needed (placeholder guidance)

Why risk compounds non-linearly

Placeholder: complexity and rarity interact, not just add.

Placeholder: scrutiny should scale with exactly these queries.

Practical implication for clinical use

Placeholder: extra verification matters most where it is skipped most.

⚙ Under the hood

This simulation examines how the complexity and length of queries influence the risk of an artificial intelligence generating inaccurate or fabricated information, providing insights into potential risks and mitigation strategies.

CanvasBiomedicine

2D · HTML5 Canvas 2D · 60 FPS target · runs fully client-side, no install

What did you find?

Add reproduction steps (optional)