todo2code

Participant: Codex (AI agent)

Understanding

Ticket-004 proved that multilingual similarity is useful for ordering candidates but unsafe as relation evidence. The next candidate therefore separates recall from acceptance: retrieval finds a small shortlist, while an audited reranker must explain an accepted module using repository-owned evidence or abstain.

Before introducing another semantic stage, the current communication boundary must be measured. The governance standard names participants through user-<identity> and ai-<provider> files; those records must remain distinct from ticket specifications and must produce an actionable response owner when human and agent intent diverge.

Execution plan

  1. Audit user-*/ai-* extraction and communication analysis on current todo2code tickets.
  2. Add red regressions for participant filename recognition, evidence-file exclusion and response ownership.
  3. Implement the minimal deterministic communication correction.
  4. Re-run the corrected analysis on todo2code and external tracked projects.
  5. Specify the candidate, decision, provenance and abstention contracts.
  6. Add red contract tests and cross-language gold projection fixtures.
  7. Implement the optional orchestration boundary outside the deterministic linker.
  8. Evaluate a constrained reranker on the six gold positives and negatives.
  9. Run tracked A/B on todo2code, subactor/platform and one additional repository selected from the existing seven-repository corpus.
  10. Manually review every newly proposed relation.
  11. Retain the implementation only if every precision and coverage criterion passes; otherwise remove it and retain the evidence.
  12. Run the full release validation and update readiness documentation.

Planned code locations

Risks

Guardrails

Actual changes

Blockers

Conclusion

Retain the communication correction and offline evidence contracts. Reject the live semantic production path until a provider-pinned candidate passes the same real-repository boundary.