[{"slug":"strategic-technical-debt","kind":"preprint","title":"Strategic Technical Debt","subtitle":"A Real Options Approach to Early-Stage Software Experimentation","authors":[{"name":"Rashid Azarang","affiliation":"Independent Researcher, San Pedro Garza García, Nuevo León, Mexico","orcid":"0009-0008-5528-4246"},{"name":"Mohammad Reza Azarang Esfandiari","affiliation":"Tecnológico de Monterrey, Monterrey, Nuevo León, Mexico","orcid":"0009-0005-7413-7610"}],"venue":"Preprint · arXiv (cs.SE)","date":"2026-08-17","question":"What is today’s marginal economic value of taking on one more unit of technical debt, given the probability that the venture survives long enough to have to repay it?","abstract":"Technical debt is treated, almost universally, as an engineering pathology: a liability incurred through haste and repaid through suffering. This paper argues that under the conditions that define early-stage software work (high hypothesis uncertainty, cheap experiments, and the freedom to abandon), deliberately incurred technical debt is not a pathology but a rationally priced financial instrument: a call option on the validated product, purchased at a discount that is largest exactly when uncertainty is highest. We make three contributions. First, a demarcation: debt is strategic when its expected cost loads on the success branch of the venture (repaid only if the hypothesis validates) and toxic when it imposes unconditional cost while held (security exposure, data loss, corrupted experimental signal), a boundary we state formally and that renders the popular \"prudent vs. reckless\" intuition testable. Second, a sequential model: a finite-horizon dynamic program over belief and debt stock whose solution yields four results: a shadow price of debt below one, equal to the risk-discounted probability of repayment; a technical-debt overhang (the belief threshold for scaling rises with the debt stock, an exercise-threshold comparative static in the tradition of McDonald and Siegel (1986), related to but mechanistically distinct from Myers (1977)); a refactoring-pivot theorem (absent carrying costs, optimal repayment concentrates at the commitment boundary, predicting a refactoring burst at product–market fit, a pattern practitioners report but, to our knowledge, no repository study has measured, registered here as a falsifiable prediction); and a volatility result (the debt build holds a call where the robust build holds the underlying, so under risk-neutral valuation mean-preserving spreads in outcome value favor debt). A pivot-salvage correction shows the folk rule \"maximum debt at maximum uncertainty\" is wrong whenever failure redirects rather than terminates the venture and the salvage differential clears the discounted cost premium. Third, a set of falsifiable propositions with a two-test primary empirical program advanced for pre-registration and execution: validation-event refactoring timing with a funding-confound design, and the first repository-history measurement of pivot salvage, with three further propositions specified as successors. Quoted quantities from the model are illustrative calibration, not estimates.","keywords":["technical debt","real options","software experimentation","startups","lean methodology","debt overhang","refactoring","pre-registration"],"status":"Under review, Journal of Systems and Software","url":"https://arxiv.org/abs/2608.16112","arxiv":"2608.16112","pdf":"https://arxiv.org/pdf/2608.16112","preregistrations":[{"id":"EP3′","title":"Refactoring timing after validation","label":"Do startups refactor right after their hypothesis is validated, or on a schedule? Measured on repository histories, with funding controlled for.","doi":"10.17605/OSF.IO/BS3CR","url":"https://osf.io/bs3cr"},{"id":"EP-Π","title":"Pivot salvage","label":"When a startup pivots, how much of the code it already wrote survives into the new direction?","url":"https://osf.io/rvx5t"}],"cover":"/research/strategic-technical-debt/cover.gif","license":"arXiv non-exclusive license","format":"external"},{"slug":"agent-memory-allocation","kind":"preprint","title":"Reading More, Finding Less","subtitle":"A Pre-Registered Anatomy of Progressive Disclosure for AI Agents","authors":[{"name":"Rashid Azarang","affiliation":"Independent Researcher, San Pedro Garza García, Nuevo León, Mexico","orcid":"0009-0008-5528-4246"}],"venue":"Preprint · Zenodo","date":"2026-08-15","question":"If you give an AI agent a hand-written index of what to read instead of a search tool, does it find the right document more often? No: it reads more files and lands on the wrong one more often.","abstract":"This paper measures how well a hand-authored digest index routes an AI agent to the right information, on a prose document corpus. Gold-reach instruments exist across the neighboring literatures; what this study contributes is the document-corpus instantiation with an attribution its neighbors do not carry: an authored one-line-digest index as the curated policy’s only way to locate documents, per-question read-level telemetry on that arm, a decomposition of failure, and a difficulty-matched oracle control. The index routed the agent to the correct document on 52.0% of 102 questions against grep’s 80.4%; of the curated policy’s 47 wrong-stops, 43 had read a wrong file. Where the index did route correctly, accuracy was 86.8%, and a difficulty-matched control shows the deficit is localization, not comprehension (Fisher p = 0.80): a policy that reads more and finds less is being mis-routed by its index. Downstream, search-then-read beat curated disclosure on accuracy (72.5% vs 47.1%). A pre-registered public replication over 141 releasable documents and 120 frozen questions, under deliberately harder criteria, returned search +12.5 pp on accuracy, wrong-stop rates of 34.2% vs 18.3% under a symmetric rule, and localization of 75.0% vs 62.5%; its formal verdict is revised because one hardened token-headroom prediction failed, reported at the same prominence as the passes. In the same program, files promoted into a durable memory directory were later read in 2 of 157 eligible cases (refuted). Predictions were frozen and analyzers committed before any data were read; the replication ships its corpus, questions, harness and adjudicator for byte-identical re-running. The staked prediction is substrate-scoped: an authored one-line-digest index over a prose corpus, used as a sole locator, routes worse than search.","keywords":["agent retrieval","retrieval routing","context engineering","LLM agents","progressive disclosure","pre-registration","replication"],"url":"https://rashidazarang.com/research/agent-memory-allocation","doi":"10.5281/zenodo.21960138","doiUrl":"https://doi.org/10.5281/zenodo.21960138","pdf":"https://rashidazarang.com/research/agent-memory-allocation/agent-memory-allocation.pdf","repository":"https://github.com/mentu-ai/agent-memory-allocation","reproducibility":"https://zenodo.org/records/21960138","cover":"/research/agent-memory-allocation/cover.gif","license":"CC BY 4.0","format":"html"},{"slug":"from-traceability-to-justifiability","kind":"preprint","title":"From Traceability to Justifiability","subtitle":"Accountability Structures in Agentic Software Engineering","authors":[{"name":"Rashid Azarang","affiliation":"Independent Researcher, San Pedro Garza García, Nuevo León, Mexico","orcid":"0009-0008-5528-4246"}],"venue":"Preprint · arXiv (cs.SE)","date":"2026-08-21","question":"When a pipeline promotes an AI system, can its records even say that the thing tested is the thing deployed? Measured across 47 platforms and 30 public repositories: not yet. Nothing emits a checkable identity of the deployed behavior by default.","abstract":"A pipeline promoting an AI system publishes records claiming the thing evaluated is the thing deployed and that the evidence licensed the transition. We measure, from public material only, whether those records can express that claim and whether it holds where declared. First, a two-class documentation survey of 47 delivery platforms (20 CI/CD, 27 model-serving/agent) under one fixed three-label protocol, graded twice (second pass blind), every consulted page pinned by content hash and date. Across 188 double-graded cells we found no platform whose default record emits a content-addressed identity of the behavioral tuple (model version, instructions, tool definitions, retrieval and runtime configuration); the blind pass grades that column default on zero of 47. Immutable nominal versioning is meanwhile arriving as the agent platforms’ default answer (16 of 27): version integers behind mutable pointers, a layer the artifact supply chain already found insufficient. Second, an instrument computes realized assurance depth from a pipeline’s published exhaust alone and compares it with the declared depth. Applied to a frozen two-stratum frame of 30 public repositories graded twice from a hashed archive, the sharpest result is a verifiability hole: seven of the 15 repositories chosen for adopting attestation tooling publish source-only releases, so the binding their workflows declare cannot be checked where declared. Where checkable it mostly checks out: five of seven measurable adopters realize the binding end to end; both shortfalls fall at identity binding. Together the results locate the field’s records structurally short of justifiability, the one rung that can refuse a transition. The survey carries an expiry clock; we state what would falsify each finding.","keywords":["software supply chain","attestation","traceability","agentic software engineering","accountability","CI/CD"],"url":"https://arxiv.org/abs/2608.23610","arxiv":"2608.23610","pdf":"https://arxiv.org/pdf/2608.23610","repository":"https://github.com/mentu-ai/from-traceability-to-justifiability","preregistrations":[{"id":"D3JFV","title":"Assurance continuity, second-site study","label":"Does the gap between declared and realized assurance reappear at a second deployment site? Predictions frozen before observation; deliberately not reported in the paper.","doi":"10.17605/OSF.IO/D3JFV","url":"https://osf.io/d3jfv"}],"license":"arXiv non-exclusive license","format":"external"},{"slug":"agent-graph-runtime","kind":"preprint","title":"The Agent Graph Runtime","subtitle":"A Unified Model for Static, Dynamic, and Hybrid Execution","authors":[{"name":"Rashid Azarang","affiliation":"Independent Researcher, San Pedro Garza García, Nuevo León, Mexico","orcid":"0009-0008-5528-4246"}],"venue":"Preprint · GitHub","date":"2026-08-22","question":"Agent systems get sorted into static workflows, dynamic planners, or hybrids. Those labels mix up four separate questions: when the graph is built, where it becomes fixed and addressable, who authors it, and what runs it. Answer them separately and the three kinds turn out to be points on one map.","abstract":"Agent systems are often classified as static workflows, dynamic planners, or hybrid arrangements. Those labels conflate when a graph is constructed, where it becomes immutable and addressable, how it is authored, and what executes it. This paper defines an agent graph runtime model in which those concerns are separately specified lifecycle coordinates over a constrained feasible region and a shared executable representation and execution substrate. Static, dynamic one-shot, staged-direct, and staged-scaffold configurations are reference points in that region, not necessarily distinct runtimes. The model separates construction, lowering, qualification, freezing, requalification, admission, execution, and evidence, and separates properties that are easily confused: executable graph identity, execution environment identity, output reproducibility, plan adequacy, and outcome. Released v1.0 on 2026-08-22 with the full six-epoch external review provenance preserved rather than summarized away.","keywords":["agent runtime","execution model","workflow graphs","reproducibility","provenance"],"url":"https://github.com/mentu-ai/agent-graph-runtime","pdf":"https://raw.githubusercontent.com/mentu-ai/agent-graph-runtime/main/paper-v1.0-preprint.pdf","repository":"https://github.com/mentu-ai/agent-graph-runtime","license":"CC BY 4.0 (documents), MIT (code)","format":"external"},{"slug":"structural-waste","kind":"under-review","title":"Structural Waste in Digital Operations","subtitle":"A Lean Theory of How the Weakest Layer Limits Capability","authors":[{"name":"Rashid Azarang","affiliation":"Independent Researcher, San Pedro Garza García, Nuevo León, Mexico","orcid":"0009-0008-5528-4246"}],"venue":"Under review","date":"2026-08-01","status":"Under review, European Journal of Information Systems","question":"Lean manufacturing made waste visible and removable. What is the equivalent inside the software that runs an operation, and is a company’s capability capped by its weakest layer? A theory, two measurement instruments, and a pre-registered 500-repository test whose frozen criterion was not met, reported in full.","abstract":"Lean production made waste visible and eliminable, yet its migration to information systems has addressed service delivery, not the architecture that determines operational capability. This paper advances a theory of structural waste: operational inefficiency from architectural misalignment between system components, distinct from technical debt and process waste. By disciplined analogy with the Toyota Production System’s relational structure it derives a five-layer modal architecture (Data, Logic, Interface, Orchestration, Feedback), maps the seven wastes to digital equivalents, and specifies three sub-dimensions: coordination overhead, semantic drift, dependency concentration. The framework yields two instruments and three propositions: separation lowers waste; flow precedes automation; capability is bounded by the least mature layer. The third is proved as a weakest-layer band with an estimable compensation budget, shown by simulation to be separable from additive models, and probed by a pre-registered 500-repository study whose frozen criterion was not met, a verdict on the proxy rather than the cross-layer claim, reported in full.","keywords":["lean","information systems","structural waste","operational capability","architecture","pre-registration"],"license":"All rights reserved while under review","format":"external"},{"slug":"finding-more-fusing-less","kind":"companion","title":"Finding More, Fusing Less","subtitle":"A Pre-Registered Locator Bake-off for AI Agents","authors":[{"name":"Rashid Azarang","affiliation":"Independent Researcher, San Pedro Garza García, Nuevo León, Mexico","orcid":"0009-0008-5528-4246"}],"venue":"Companion preprint · Zenodo","date":"2026-08-16","question":"Search still missed one document in five. Was that the limit of search, or just bad ranking? Mostly ranking: a better ranker found far more, and combining search methods made things worse.","abstract":"The parent study’s winning arm still missed a fifth of the corpus. This pre-registered bake-off shows most of that ceiling was a ranking problem (BM25 +21.7 pp on localization) and retires the shipped fusion default, which failed its own frozen prediction. Verdict: revised, both outcomes at equal prominence; independently validated before writing.","keywords":["agent retrieval","BM25","rank fusion","locator","pre-registration"],"url":"https://doi.org/10.5281/zenodo.21969901","doi":"10.5281/zenodo.21969901","doiUrl":"https://doi.org/10.5281/zenodo.21969901","reproducibility":"https://zenodo.org/records/21969901","license":"CC BY 4.0","format":"external"},{"slug":"mentu-navigator","kind":"software","title":"mentu-navigator","question":"The tool the two retrieval studies produced: a way for AI agents to find and read exactly the right part of a codebase, where every default was chosen by an experiment rather than by taste.","description":"Read-only, provenance-first repository navigation for humans and AI agents: a CLI (mentu-nav) and MCP server exposing a pinned four-primitive contract (locate, read_range, open, handles). The locator composes an in-memory Okapi BM25 index with per-language Snowball analyzers and a hardened exact-search leg. Every performance-relevant default traces to a registered, mechanically adjudicated study; one default has already been retired by its frozen prediction.","date":"2026-08-22","doi":"10.5281/zenodo.22061819","doiUrl":"https://doi.org/10.5281/zenodo.22061819","repository":"https://github.com/mentu-ai/mentu-navigator","supplementTo":["10.5281/zenodo.21960138","10.5281/zenodo.21969901"],"license":"Open source"},{"slug":"agent-harness","kind":"working-paper","status":"Working paper · 2026 · confirmatory result pending","title":"Agent Harness: Deterministic Control for Probabilistic Systems","subtitle":"A five-plane decomposition and a conditional invariant-preservation result","question":"Can you get a hard guarantee out of a system built on a model that is not deterministic? Yes, if the model can propose anything but commit nothing without passing a gate you control.","description":"Separates proposal authority from effect authority in an agent harness, then proves that under complete mediation of a declared effect boundary and a sound commit gate, every reachable protected state preserves the declared invariant, regardless of what the model proposes. The corollary is the useful part: model stochasticity stops mattering to the safety argument once proposals carry no independent commit authority. Study pre-registered before analyzer code was written."},{"slug":"return-base-rate","kind":"working-paper","status":"Working paper · 2026 · preprint pending release","title":"The Return Base Rate","subtitle":"Captured knowledge is almost never reused in a production AI-agent system","question":"AI agents save notes for later all the time. In a real production system, how often does anything saved get read again? Almost never, and this paper measures where in the pipeline it gets lost.","description":"The agent-memory literature measures how well agents use prior context on evaluations where the answer-bearing memory is guaranteed present. This measures the step upstream: in a live system, does captured knowledge get surfaced and reused at all? Over an append-only, hash-chained ledger, return is effectively zero. Decomposes that into a funnel, localizes the stage the instrument could not previously see, and pre-registers the intervention."}]