Research

Papers with their evidence attached: pre-registered experiments, published data, and results anyone can re-run.

Preprint · arXiv:2608.16112 [cs.SE]Strategic Technical DebtWhat is today’s marginal economic value of taking on one more unit of technical debt, given the probability that the venture survives long enough to have to repay it?Under review, Journal of Systems and SoftwareRead on arXiv →
Strategic Technical Debt: A Real Options Approach to Early-Stage Software Experimentation
Preprint · DOI 10.5281/zenodo.21960138Reading More, Finding LessIf you give an AI agent a hand-written index of what to read instead of a search tool, does it find the right document more often? No: it reads more files and lands on the wrong one more often.Read paper →
Reading More, Finding Less: A Pre-Registered Anatomy of Progressive Disclosure for AI Agents

Public preprints on arXiv and Zenodo, with their replication packages where there is data, and manuscripts that are not yet public.

Preprint

From Traceability to Justifiability

Accountability Structures in Agentic Software Engineering

When a pipeline promotes an AI system, can its records even say that the thing tested is the thing deployed? Measured across 47 platforms and 30 public repositories: not yet. Nothing emits a checkable identity of the deployed behavior by default.

Preprint

The Agent Graph Runtime

A Unified Model for Static, Dynamic, and Hybrid Execution

Agent systems get sorted into static workflows, dynamic planners, or hybrids. Those labels mix up four separate questions: when the graph is built, where it becomes fixed and addressable, who authors it, and what runs it. Answer them separately and the three kinds turn out to be points on one map.

Preprint

Outbound Engineering

A Model of Unsolicited B2B Outreach as a Measured System

Cold email gets judged by vendor dashboards with no controls. What changes when outbound is built as a measured system: evidence behind every address, a record of each send before its outcome arrives, and a gate that can refuse to send? A definition and seven testable propositions; no empirical result is claimed yet.

In revision

Structural Waste in Digital Operations

A Lean Theory of How the Weakest Layer Limits Capability

Lean manufacturing made waste visible and removable. What is the equivalent inside the software that runs an operation, and is a company’s capability capped by its weakest layer? A theory, two measurement instruments, and a pre-registered 500-repository test whose frozen criterion was not met, reported in full.

Follow-ups to the papers above: the study that closed the gap the first one left open, and the tool both of them produced, archived with a DOI.

Software

mentu-navigator

The tool the two retrieval studies produced: a way for AI agents to find and read exactly the right part of a codebase, where every default was chosen by an experiment rather than by taste.

Predictions and analysis code frozen in public before any data was touched, so results are judged against what was promised, pass or fail.

Machine-readable: /api/research (JSON, every record above with abstracts and identifiers) · ORCID 0009-0008-5528-4246