Research
Papers with their evidence attached: pre-registered experiments, published data, and results anyone can re-run.
Preprints and submissions
Public preprints on arXiv and Zenodo, with their replication packages where there is data, and manuscripts that are not yet public.
PreprintAccountability Structures in Agentic Software Engineering
When a pipeline promotes an AI system, can its records even say that the thing tested is the thing deployed? Measured across 47 platforms and 30 public repositories: not yet. Nothing emits a checkable identity of the deployed behavior by default.
PreprintA Unified Model for Static, Dynamic, and Hybrid Execution
Agent systems get sorted into static workflows, dynamic planners, or hybrids. Those labels mix up four separate questions: when the graph is built, where it becomes fixed and addressable, who authors it, and what runs it. Answer them separately and the three kinds turn out to be points on one map.
PreprintA Model of Unsolicited B2B Outreach as a Measured System
Cold email gets judged by vendor dashboards with no controls. What changes when outbound is built as a measured system: evidence behind every address, a record of each send before its outcome arrives, and a gate that can refuse to send? A definition and seven testable propositions; no empirical result is claimed yet.
In revisionStructural Waste in Digital Operations
A Lean Theory of How the Weakest Layer Limits Capability
Lean manufacturing made waste visible and removable. What is the equivalent inside the software that runs an operation, and is a company’s capability capped by its weakest layer? A theory, two measurement instruments, and a pre-registered 500-repository test whose frozen criterion was not met, reported in full.
Companion studies and software
Follow-ups to the papers above: the study that closed the gap the first one left open, and the tool both of them produced, archived with a DOI.
CompanionA Pre-Registered Locator Bake-off for AI Agents
Search still missed one document in five. Was that the limit of search, or just bad ranking? Mostly ranking: a better ranker found far more, and combining search methods made things worse.
SoftwareThe tool the two retrieval studies produced: a way for AI agents to find and read exactly the right part of a codebase, where every default was chosen by an experiment rather than by taste.
Pre-registrations
Predictions and analysis code frozen in public before any data was touched, so results are judged against what was promised, pass or fail.
EP3′Pre-registered for Strategic Technical Debt
Do startups refactor right after their hypothesis is validated, or on a schedule? Measured on repository histories, with funding controlled for.
EP-ΠPre-registered for Strategic Technical Debt
When a startup pivots, how much of the code it already wrote survives into the new direction?
D3JFVPre-registered for From Traceability to Justifiability
Does the gap between declared and realized assurance reappear at a second deployment site? Predictions frozen before observation; deliberately not reported in the paper.
Machine-readable: /api/research (JSON, every record above with abstracts and identifiers) · ORCID 0009-0008-5528-4246