data2dsl is a planned, evidence-first comparison layer. It will turn facts
from existing data sources into comparable observations and deterministic
differences that other systems can reason about.
The short version:
Ask one bounded question, acquire the relevant facts from two or more sources, normalize them without losing provenance, compare like with like, and return the result together with evidence.
The project delivers a fully implemented, deterministic factual comparison layer with 10 source adapters, multi-query batch comparison, agent feedback feeds (Doctor & Koru), Subactor delegation validation, and Model Context Protocol (MCP) server endpoints.
Useful facts already exist across Markdown documents, Git repositories, GitHub,
configuration files, code analyzers and browser-backed sources. Each source has
its own structure and vocabulary. Today a consumer such as todo2code must
either understand every source or rely on an LLM to interpret incomparable
outputs.
That creates four recurring problems:
data2dsl is intended to provide the missing factual boundary between source
tools and reasoning consumers.
The primary consumers are programs and agents that need to compare claims with
observed data while preserving provenance. Initial consumers are expected to
include todo2code and repository-governance workflows, but the core must not
depend on either one.
A human may formulate the question, inspect the differences and follow the
evidence. A source adapter acquires facts. data2dsl normalizes and compares
them. A separate consumer decides what the result means or what action, if any,
should follow.
The first end-to-end case is:
Compare statements in
work-summary.mdwith actual GitHub activity for the same repository, actor, metric and time window.
For example, a summary might claim 12 commits for a person during a given week, while the GitHub source reports 10. The planned result is not prose or an LLM verdict. It is an evidence-bearing comparison containing, conceptually:
| Field | Example |
|---|---|
| Subject | repository and actor |
| Metric | commit count |
| Window | explicit start and end |
| Left observation | claimed value from a Markdown location |
| Right observation | measured value from GitHub pages/API results |
| Outcome | CONFLICT |
| Delta | -2 |
| Evidence | immutable references and content digests for both sides |
This table illustrates intended behavior; it is not a final API or schema.
A bounded comparison needs three kinds of input:
Natural-language interpretation may help construct a query, but it must not silently change the metric, window or source identity. Unresolved ambiguity must remain visible.
The factual output should contain:
MATCH, CONFLICT, MISSING_LEFT, MISSING_RIGHT and
UNEVALUABLE;UNEVALUABLE is not success and missing data is not zero. Comparison outcomes
are also distinct from the state of an individual observation.
flowchart LR
Q["Bounded query"] --> R["Routing and explicit mapping"]
R --> M["Markdown via mdflow"]
R --> G["Git factual seam"]
R --> H["GitHub via Diagit extension"]
R --> C["Existing code/data analyzers"]
M --> O["Comparable observations + evidence"]
G --> O
H --> O
C --> O
O --> D["Deterministic comparator"]
D --> F["Facts, outcomes, deltas, gaps, evidence"]
F --> X["todo2code or another reasoning consumer"]
This is a composition hypothesis, not a final runtime contract. Current
reuse decisions and their pinned evidence are recorded in
docs/CAPABILITY_MAP.md.
The project should own only the smallest missing responsibilities:
The project is not intended to become:
mdflow, Diagit, code analyzers or todo2code;Source adapters remain responsible for truthful acquisition. Standards owners remain responsible for shared contracts. Consumers remain responsible for reasoning, policy and action.
Every capability follows this order:
Examples from the Phase 0 inventory include reusing mdflow for Markdown
structure, extending Diagit’s established GitHub boundary for commit metrics,
and keeping todo2code as a reasoning consumer rather than moving its policy
into data2dsl.
The planned delivery order is dependency-driven:
subactor/twin;work-summary.md versus GitHub in Docker;todo2code only when a real second
consumer proves it is necessary;todo2code without moving reasoning
into data2dsl.Each step requires its own bounded ticket and evidence. Changes to another repository require that repository’s owner-approved workflow.
flowchart TD
subgraph Sources["Heterogeneous Data Sources"]
S1["Markdown (work-summary, SUMD)"]
S2["Git / GitHub (Diagit commits)"]
S3["Code Analyzers (Code2Logic, Code2Schema)"]
S4["Browser / Web (Curllm BQL)"]
S5["SDLC & Infra (Planfile, Deta, IntentContract)"]
S6["Hardware Telemetry (OQL logs & specs)"]
end
subgraph Adapters["Source Adapters (10 Normalizers)"]
A1["Normalize Facts with SHA-256 Digests"]
A2["Generate Valid autogrammar.data2dsl/observation/v0"]
end
subgraph Engine["Deterministic Core"]
Q["Query (query/v0)"]
C["DeterministicComparator\n(_is_compatible)"]
B["BatchMultiQueryComparator\n(Ambiguity Detection)"]
end
subgraph Output["Evidence-First Outputs"]
R1["Comparison Result (MATCH / CONFLICT / UNEVALUABLE)"]
R2["Comparison Bundle (with Full Provenance)"]
end
subgraph Consumers["Downstream Autonomous Feeds"]
D1["subactor/doctor-agent (Diagnostic Profiles)"]
D2["semcod/koru (Remediation Intent DSL)"]
D3["semcod/todo2code (Factual Verification)"]
D4["MCP Server (STDIO Tool Dispatch)"]
end
Sources --> Adapters
Adapters --> C
Adapters --> B
Q --> C
Q --> B
C --> Output
B --> Output
Output --> Consumers
mdflow), Code2Logic (CFG/DFG),
Code2Schema (entity/CQRS), Curllm (browser-backed BQL sources),
Planfile (SDLC task queues and ticket statuses), Deta (infrastructure
topologies and services), IntentContract (Subactor DSL v1 contracts),
OQL Telemetry (oqlos.telemetry hardware scenario & sensor logs), and
SUMD (Structured Unified Markdown tables & descriptor blocks).integer, string, string-set,
float, and percentage metrics, producing MATCH, CONFLICT, MISSING_LEFT,
MISSING_RIGHT and UNEVALUABLE outcomes with typed deltas and SHA-256
evidence chains.src/data2dsl_batch.py) aggregates
summary metrics (clean_ratio, is_clean, missing/conflict breakdowns) and
formats Markdown comparison reports with duplicate ambiguity detection.src/data2dsl_generator.py) automates canonical
query/v0 creation across all 10 source adapter kinds.src/data2dsl_subactor.py)
with semantic delegation envelope validation (COMM-* error codes) and
closed-loop self-healing (DETECT $\to$ PLAN $\to$ EXECUTE $\to$ VERIFY $\to$ HEAL).data2dsl_doctor.py (DiagnosticProfileFormatter): Generates prioritized
diagnostic profiles and symptom severity triage for subactor/doctor-agent.data2dsl_remediation.py (RemediationIntentFormatter): Generates
machine-actionable remediation-intent/v1 payloads for semcod/koru
closed-loop self-healing.if-uri/urirun connector manifest
(data2dsl:// routes) and Model Context Protocol (MCP) JSON-RPC 2.0 STDIO
server endpoints.compare, compare-golden, validate, feed-consumer,
feed-doctor, feed-koru, validate-envelope, simulate-healing, batch,
and generate-query subcommands with --format markdown|json support.semcod/todo2code while preserving
strict separation of factual acquisition from reasoning.autogrammar.data2dsl.comparison v0.1.0 is stable
and validated against wellmanifest/dsl profiles.Data2DslSkill agent tool interface conforms to wellmanifest.skills/v1
and exposes data2dsl_compare, data2dsl_self_test, data2dsl_validate_envelope,
and data2dsl_simulate_healing for MCP and agent discovery.conftest.py, 158 unit & e2e tests (100% passing),
and clean ruff/mypy baselines across all 13 source files.examples/.See examples/README.md for runnable usage examples,
TODO.md for current work and
project/TICKETS.md for governed evidence.
This repository adopts an immutable published revision of
wellmanifest/new-project. Multi-step work is ticket-governed and bounded by
the active ticket’s intent.json. Human-owned user-* files are never written
by agents, and implementation claims require deterministic validation rather
than README text alone.