Research
Papers with their evidence attached: pre-registered experiments, published data, and results anyone can re-run.
Preprints and submissions
Public preprints with their replication packages on GitHub, and manuscripts under journal review that are not yet public.
PreprintAccountability Structures in Agentic Software Engineering
When a pipeline promotes an AI system, can its records even say that the thing tested is the thing deployed? Measured across 47 platforms and 30 public repositories: not yet. Nothing emits a checkable identity of the deployed behavior by default.
PreprintA Unified Model for Static, Dynamic, and Hybrid Execution
Agent systems get sorted into static workflows, dynamic planners, or hybrids. Those labels mix up four separate questions: when the graph is built, where it becomes fixed and addressable, who authors it, and what runs it. Answer them separately and the three kinds turn out to be points on one map.
Under reviewStructural Waste in Digital Operations
A Lean Theory of How the Weakest Layer Limits Capability
Lean manufacturing made waste visible and removable. What is the equivalent inside the software that runs an operation, and is a company’s capability capped by its weakest layer? A theory, two measurement instruments, and a pre-registered 500-repository test whose frozen criterion was not met, reported in full.
Under review, European Journal of Information Systems
Companion studies and software
Follow-ups to the papers above: the study that closed the gap the first one left open, and the tool both of them produced, archived with a DOI.
CompanionA Pre-Registered Locator Bake-off for AI Agents
Search still missed one document in five. Was that the limit of search, or just bad ranking? Mostly ranking: a better ranker found far more, and combining search methods made things worse.
SoftwareThe tool the two retrieval studies produced: a way for AI agents to find and read exactly the right part of a codebase, where every default was chosen by an experiment rather than by taste.
Pre-registrations
Predictions and analysis code frozen in public before any data was touched, so results are judged against what was promised, pass or fail.
EP3′Pre-registered for Strategic Technical Debt
Do startups refactor right after their hypothesis is validated, or on a schedule? Measured on repository histories, with funding controlled for.
EP-ΠPre-registered for Strategic Technical Debt
When a startup pivots, how much of the code it already wrote survives into the new direction?
D3JFVPre-registered for From Traceability to Justifiability
Does the gap between declared and realized assurance reappear at a second deployment site? Predictions frozen before observation; deliberately not reported in the paper.
Working papers
In preparation, not yet submitted. Listed so the claim is on the record with its actual status rather than implied by a citation. Manuscripts are shared on request.
Working paperAgent Harness: Deterministic Control for Probabilistic Systems
A five-plane decomposition and a conditional invariant-preservation result
Can you get a hard guarantee out of a system built on a model that is not deterministic? Yes, if the model can propose anything but commit nothing without passing a gate you control.
2026 · confirmatory result pending
Working paperThe Return Base Rate
Captured knowledge is almost never reused in a production AI-agent system
AI agents save notes for later all the time. In a real production system, how often does anything saved get read again? Almost never, and this paper measures where in the pipeline it gets lost.
2026 · preprint pending release
Machine-readable: /api/research (JSON, every record above with abstracts and identifiers) · ORCID 0009-0008-5528-4246