DISTRIBUTED LLM OBSERVATORY

Build agents.
Test agents.
Observe AI systems.

An open-source, evidence-first framework for designing agent configurations, evaluating real agent behavior, and comparing reproducible AI observations across time and regions.

Browser workflows run locally through the DLLO Agent Lab bridge.

v0.1.2 PyPI Release

From agent design to reproducible observation.

DLLO keeps requirements, evidence, evaluation, provenance and interpretation explicit throughout the workflow.

Agent Starter visual

Build your agent.

Turn goals, constraints, hardware evidence and explicit preferences into an evidence-backed architecture and concrete stack recommendation.

  • Adaptive questioning
  • Five supported agent goals
  • Hardware and runtime evidence
  • Why / Why Not explanations
  • No hidden global winner
Explore Agent Starter →
Agent evaluation visual

Evaluate real behavior.

Connect an agent through the Local SUT Protocol and evaluate observable behavior using observer-owned expectations and evidence.

  • Agent Protocol Core 1.0
  • Tool and action validation
  • Error recovery and branching
  • Persistent run artifacts
  • Technical reports
Test your agent →
DLLO Observatory report visual

Observe what changed.

Compare sufficiently compatible observations across time and geographic regions while preserving provenance and rejected comparisons.

  • Temporal comparison
  • Geographic comparison
  • Observation-pair discovery
  • Explicit comparability rules
  • No unsupported causal claims
Explore the Observatory →

Conservative by design.

UNKNOWN ≠ FAILURE

Missing evidence remains unknown instead of becoming a negative claim.

NO HIDDEN RANKING

DLLO does not invent a global winner when the evidence supports multiple valid outcomes.

OBSERVER ≠ SUT

The system under test does not own the expectations used to certify its behavior.

PROVENANCE FIRST

Observations preserve what was measured, where, when and under which benchmark context.

Validated from a fresh clone.

1552 automated tests
1434 Python tests
118 browser tests
v0.1.2 PyPI Release