[{"slug":"strategic-technical-debt","kind":"preprint","title":"Strategic Technical Debt","subtitle":"A Real Options Approach to Early-Stage Software Experimentation","authors":[{"name":"Rashid Azarang","affiliation":"Independent Researcher, San Pedro Garza García, Nuevo León, Mexico","orcid":"0009-0008-5528-4246"},{"name":"Mohammad Reza Azarang Esfandiari","affiliation":"Tecnológico de Monterrey, Monterrey, Nuevo León, Mexico","orcid":"0009-0005-7413-7610"}],"venue":"Preprint · arXiv (cs.SE)","date":"2026-08-17","question":"What is today’s marginal economic value of taking on one more unit of technical debt, given the probability that the venture survives long enough to have to repay it?","abstract":"Technical debt is treated, almost universally, as an engineering pathology: a liability incurred through haste and repaid through suffering. This paper argues that under the conditions that define early-stage software work (high hypothesis uncertainty, cheap experiments, and the freedom to abandon), deliberately incurred technical debt is not a pathology but a rationally priced financial instrument: a call option on the validated product, purchased at a discount that is largest exactly when uncertainty is highest. We make three contributions. First, a demarcation: debt is strategic when its expected cost loads on the success branch of the venture (repaid only if the hypothesis validates) and toxic when it imposes unconditional cost while held (security exposure, data loss, corrupted experimental signal), a boundary we state formally and that renders the popular \"prudent vs. reckless\" intuition testable. Second, a sequential model: a finite-horizon dynamic program over belief and debt stock whose solution yields four results: a shadow price of debt below one, equal to the risk-discounted probability of repayment; a technical-debt overhang (the belief threshold for scaling rises with the debt stock, an exercise-threshold comparative static in the tradition of McDonald and Siegel (1986), related to but mechanistically distinct from Myers (1977)); a refactoring-pivot theorem (absent carrying costs, optimal repayment concentrates at the commitment boundary, predicting a refactoring burst at product–market fit, a pattern practitioners report but, to our knowledge, no repository study has measured, registered here as a falsifiable prediction); and a volatility result (the debt build holds a call where the robust build holds the underlying, so under risk-neutral valuation mean-preserving spreads in outcome value favor debt). A pivot-salvage correction shows the folk rule \"maximum debt at maximum uncertainty\" is wrong whenever failure redirects rather than terminates the venture and the salvage differential clears the discounted cost premium. Third, a set of falsifiable propositions with a two-test primary empirical program advanced for pre-registration and execution: validation-event refactoring timing with a funding-confound design, and the first repository-history measurement of pivot salvage, with three further propositions specified as successors. Quoted quantities from the model are illustrative calibration, not estimates.","keywords":["technical debt","real options","software experimentation","startups","lean methodology","debt overhang","refactoring","pre-registration"],"status":"Under review, Journal of Systems and Software","url":"https://arxiv.org/abs/2608.16112","arxiv":"2608.16112","arxivCategory":"cs.SE","pdf":"https://arxiv.org/pdf/2608.16112","preregistrations":[{"id":"EP3′","title":"Refactoring timing after validation","label":"Do startups refactor right after their hypothesis is validated, or on a schedule? Measured on repository histories, with funding controlled for.","doi":"10.17605/OSF.IO/BS3CR","url":"https://osf.io/bs3cr","addendum":"https://osf.io/nmcpw"},{"id":"EP-Π","title":"Pivot salvage","label":"When a startup pivots, how much of the code it already wrote survives into the new direction?","doi":"10.17605/OSF.IO/RVX5T","url":"https://osf.io/rvx5t","addendum":"https://osf.io/nmcpw"}],"cover":"/research/strategic-technical-debt/cover.gif","license":"arXiv non-exclusive license","format":"external"},{"slug":"agent-memory-allocation","kind":"preprint","title":"Reading More, Finding Less","subtitle":"A Pre-Registered Anatomy of Progressive Disclosure for AI Agents","authors":[{"name":"Rashid Azarang","affiliation":"Independent Researcher, San Pedro Garza García, Nuevo León, Mexico","orcid":"0009-0008-5528-4246"}],"venue":"Preprint · Zenodo","date":"2026-08-15","question":"If you give an AI agent a hand-written index of what to read instead of a search tool, does it find the right document more often? No: it reads more files and lands on the wrong one more often.","abstract":"This paper measures how well a hand-authored digest index routes an AI agent to the right information, on a prose document corpus. Gold-reach instruments exist across the neighboring literatures; what this study contributes is the document-corpus instantiation with an attribution its neighbors do not carry: an authored one-line-digest index as the curated policy’s only way to locate documents, per-question read-level telemetry on that arm, a decomposition of failure, and a difficulty-matched oracle control. The index routed the agent to the correct document on 52.0% of 102 questions against grep’s 80.4%; of the curated policy’s 47 wrong-stops, 43 had read a wrong file. Where the index did route correctly, accuracy was 86.8%, and a difficulty-matched control shows the deficit is localization, not comprehension (Fisher p = 0.80): a policy that reads more and finds less is being mis-routed by its index. Downstream, search-then-read beat curated disclosure on accuracy (72.5% vs 47.1%). A pre-registered public replication over 141 releasable documents and 120 frozen questions, under deliberately harder criteria, returned search +12.5 pp on accuracy, wrong-stop rates of 34.2% vs 18.3% under a symmetric rule, and localization of 75.0% vs 62.5%; its formal verdict is revised because one hardened token-headroom prediction failed, reported at the same prominence as the passes. In the same program, files promoted into a durable memory directory were later read in 2 of 157 eligible cases (refuted). Predictions were frozen and analyzers committed before any data were read; the replication ships its corpus, questions, harness and adjudicator for byte-identical re-running. The staked prediction is substrate-scoped: an authored one-line-digest index over a prose corpus, used as a sole locator, routes worse than search.","keywords":["agent retrieval","retrieval routing","context engineering","LLM agents","progressive disclosure","pre-registration","replication"],"url":"https://rashidazarang.com/research/agent-memory-allocation","doi":"10.5281/zenodo.21960138","doiUrl":"https://doi.org/10.5281/zenodo.21960138","pdf":"https://rashidazarang.com/research/agent-memory-allocation/agent-memory-allocation.pdf","repository":"https://github.com/mentu-ai/agent-memory-allocation","reproducibility":"https://zenodo.org/records/21960138","cover":"/research/agent-memory-allocation/cover.gif","license":"CC BY 4.0","format":"html"},{"slug":"from-traceability-to-justifiability","kind":"preprint","title":"From Traceability to Justifiability","subtitle":"Accountability Structures in Agentic Software Engineering","authors":[{"name":"Rashid Azarang","affiliation":"Independent Researcher, San Pedro Garza García, Nuevo León, Mexico","orcid":"0009-0008-5528-4246"}],"venue":"Preprint · arXiv (cs.SE)","date":"2026-08-21","question":"When a pipeline promotes an AI system, can its records even say that the thing tested is the thing deployed? Measured across 47 platforms and 30 public repositories: not yet. Nothing emits a checkable identity of the deployed behavior by default.","abstract":"A pipeline promoting an AI system publishes records claiming the thing evaluated is the thing deployed and that the evidence licensed the transition. We measure, from public material only, whether those records can express that claim and whether it holds where declared. First, a two-class documentation survey of 47 delivery platforms (20 CI/CD, 27 model-serving/agent) under one fixed three-label protocol, graded twice (second pass blind), every consulted page pinned by content hash and date. Across 188 double-graded cells we found no platform whose default record emits a content-addressed identity of the behavioral tuple (model version, instructions, tool definitions, retrieval and runtime configuration); the blind pass grades that column default on zero of 47. Immutable nominal versioning is meanwhile arriving as the agent platforms’ default answer (16 of 27): version integers behind mutable pointers, a layer the artifact supply chain already found insufficient. Second, an instrument computes realized assurance depth from a pipeline’s published exhaust alone and compares it with the declared depth. Applied to a frozen two-stratum frame of 30 public repositories graded twice from a hashed archive, the sharpest result is a verifiability hole: seven of the 15 repositories chosen for adopting attestation tooling publish source-only releases, so the binding their workflows declare cannot be checked where declared. Where checkable it mostly checks out: five of seven measurable adopters realize the binding end to end; both shortfalls fall at identity binding. Together the results locate the field’s records structurally short of justifiability, the one rung that can refuse a transition. The survey carries an expiry clock; we state what would falsify each finding.","keywords":["software supply chain","attestation","traceability","agentic software engineering","accountability","CI/CD"],"url":"https://arxiv.org/abs/2608.23610","arxiv":"2608.23610","arxivCategory":"cs.SE","pdf":"https://arxiv.org/pdf/2608.23610","repository":"https://github.com/mentu-ai/from-traceability-to-justifiability","preregistrations":[{"id":"D3JFV","title":"Assurance continuity, second-site study","label":"Does the gap between declared and realized assurance reappear at a second deployment site? Predictions frozen before observation; deliberately not reported in the paper.","doi":"10.17605/OSF.IO/D3JFV","url":"https://osf.io/d3jfv"}],"license":"arXiv non-exclusive license","format":"external"},{"slug":"agent-graph-runtime","kind":"preprint","title":"The Agent Graph Runtime","subtitle":"A Unified Model for Static, Dynamic, and Hybrid Execution","authors":[{"name":"Rashid Azarang","affiliation":"Independent Researcher, San Pedro Garza García, Nuevo León, Mexico","orcid":"0009-0008-5528-4246"}],"venue":"Preprint · Zenodo","date":"2026-08-22","question":"Agent systems get sorted into static workflows, dynamic planners, or hybrids. Those labels mix up four separate questions: when the graph is built, where it becomes fixed and addressable, who authors it, and what runs it. Answer them separately and the three kinds turn out to be points on one map.","abstract":"Agent systems are often classified as static workflows, dynamic planners, or hybrid arrangements. Those labels conflate when a graph is constructed, where it becomes immutable and addressable, how it is authored, and what executes it. This paper defines an agent graph runtime model in which those concerns are separately specified lifecycle coordinates over a constrained feasible region and a shared executable representation and execution substrate. Static, dynamic one-shot, staged-direct, and staged-scaffold configurations are reference points in that region, not necessarily distinct runtimes. The model separates construction, lowering, qualification, freezing, requalification, admission, execution, and evidence, and separates properties that are easily confused: executable graph identity, execution environment identity, output reproducibility, plan adequacy, and outcome. Released v1.0 on 2026-08-22 and archived on Zenodo as v1.0.1, build-corrected on 2026-08-26 so all 45 references render, together with the reproducibility supplement, the frozen build projection, and the full provenance of six attempted external certification epochs, the release-gate refusals included as negative evidence. The paper reports a mechanical invariant exercise on a frozen historical pilot and certifies its own novelty as uncertified; comparative superiority and generality claims are excluded by design.","keywords":["agent runtime","execution model","workflow graphs","software lifecycle","reproducibility","provenance"],"url":"https://doi.org/10.5281/zenodo.22118741","doi":"10.5281/zenodo.22118741","doiUrl":"https://doi.org/10.5281/zenodo.22118741","pdf":"https://raw.githubusercontent.com/mentu-ai/agent-graph-runtime/main/paper-v1.0-preprint.pdf","repository":"https://github.com/mentu-ai/agent-graph-runtime","reproducibility":"https://zenodo.org/records/22118742","license":"CC BY 4.0 (documents), MIT (code)","format":"external"},{"slug":"outbound-engineering","kind":"preprint","title":"Outbound Engineering","subtitle":"A Model of Unsolicited B2B Outreach as a Measured System","authors":[{"name":"Rashid Azarang","affiliation":"Independent Researcher, San Pedro Garza García, Nuevo León, Mexico","orcid":"0009-0008-5528-4246"}],"venue":"Preprint · Zenodo","date":"2026-09-11","question":"Cold email gets judged by vendor dashboards with no controls. What changes when outbound is built as a measured system: evidence behind every address, a record of each send before its outcome arrives, and a gate that can refuse to send? A definition and seven testable propositions; no empirical result is claimed yet.","abstract":"Unsolicited business-to-business outreach by email is practised as writing and volume, and evaluated through vendor aggregates without controls. The sales-technology literature studies the adoption and use of tools by salespeople; the deliverability standards instrument the receiving side of a message’s fate. Neither treats the operator’s sending system as an artifact with acceptance criteria. This paper proposes outbound engineering, defined as building and running unsolicited B2B outreach as a measured system, and specifies it by three membership properties: provenance of every address, a message-level record that carries each send’s treatment before its outcome arrives, and a gate that can refuse a send. Three further properties present in every instance the author has built, namely isolation, reading-driven volume and behaviour-driven sequencing, are held as design hypotheses rather than as definitional. The construct is positioned against eight neighbouring constructs, feedback control is adopted as the method theory, and seven propositions are stated with their outcome variables, floors and confounds. The contribution is a treatment schema for a domain whose outcome schema has been standardised for two decades, in delivery status notifications (RFC 3464) and feedback reports (RFC 5965), and whose treatment side has not. No empirical result is claimed; a bounded pre-registration plan is attached as an appendix. Declaration of interest: the author sells a managed outbound service under the name proposed in the paper.","keywords":["outbound prospecting","cold email","sales technology","sender reputation","data provenance","pre-registration","conceptual article"],"url":"https://doi.org/10.5281/zenodo.22698746","doi":"10.5281/zenodo.22698746","doiUrl":"https://doi.org/10.5281/zenodo.22698746","pdf":"https://zenodo.org/records/22698747/files/outbound-engineering.pdf","reproducibility":"https://zenodo.org/records/22698747","essay":"https://rashidazarang.com/c/outbound-engineering","license":"CC BY 4.0","format":"external"},{"slug":"structural-waste","kind":"in-revision","title":"Structural Waste in Digital Operations","subtitle":"A Lean Theory of How the Weakest Layer Limits Capability","authors":[{"name":"Rashid Azarang","affiliation":"Independent Researcher, San Pedro Garza García, Nuevo León, Mexico","orcid":"0009-0008-5528-4246"},{"name":"Mohammad Reza Azarang Esfandiari","affiliation":"Tecnológico de Monterrey, Monterrey, Nuevo León, Mexico","orcid":"0009-0005-7413-7610"}],"venue":"In revision","date":"2026-08-01","question":"Lean manufacturing made waste visible and removable. What is the equivalent inside the software that runs an operation, and is a company’s capability capped by its weakest layer? A theory, two measurement instruments, and a pre-registered 500-repository test whose frozen criterion was not met, reported in full.","abstract":"Lean production made waste visible and eliminable, yet its migration to information systems has addressed service delivery, not the architecture that determines operational capability. This paper advances a theory of structural waste: operational inefficiency from architectural misalignment between system components, distinct from technical debt and process waste. By disciplined analogy with the Toyota Production System’s relational structure it derives a five-layer modal architecture (Data, Logic, Interface, Orchestration, Feedback), maps the seven wastes to digital equivalents, and specifies three sub-dimensions: coordination overhead, semantic drift, dependency concentration. The framework yields two instruments and three propositions: separation lowers waste; flow precedes automation; capability is bounded by the least mature layer. The third is proved as a weakest-layer band with an estimable compensation budget, shown by simulation to be separable from additive models, and probed by a pre-registered 500-repository study whose frozen criterion was not met, a verdict on the proxy rather than the cross-layer claim, reported in full.","keywords":["lean","information systems","structural waste","operational capability","architecture","pre-registration"],"license":"Manuscript not yet public","format":"external"},{"slug":"finding-more-fusing-less","kind":"companion","title":"Finding More, Fusing Less","subtitle":"A Pre-Registered Locator Bake-off for AI Agents","authors":[{"name":"Rashid Azarang","affiliation":"Independent Researcher, San Pedro Garza García, Nuevo León, Mexico","orcid":"0009-0008-5528-4246"}],"venue":"Companion preprint · Zenodo","date":"2026-08-16","question":"Search still missed one document in five. Was that the limit of search, or just bad ranking? Mostly ranking: a better ranker found far more, and combining search methods made things worse.","abstract":"The parent study’s winning arm still missed a fifth of the corpus. This pre-registered bake-off shows most of that ceiling was a ranking problem (BM25 +21.7 pp on localization) and retires the shipped fusion default, which failed its own frozen prediction. Verdict: revised, both outcomes at equal prominence; independently validated before writing.","keywords":["agent retrieval","BM25","rank fusion","locator","pre-registration"],"url":"https://doi.org/10.5281/zenodo.21969901","doi":"10.5281/zenodo.21969901","doiUrl":"https://doi.org/10.5281/zenodo.21969901","reproducibility":"https://zenodo.org/records/21969901","license":"CC BY 4.0","format":"external"},{"slug":"mentu-navigator","kind":"software","title":"mentu-navigator","question":"The tool the two retrieval studies produced: a way for AI agents to find and read exactly the right part of a codebase, where every default was chosen by an experiment rather than by taste.","description":"Read-only, provenance-first repository navigation for humans and AI agents: a CLI (mentu-nav) and MCP server exposing a pinned four-primitive contract (locate, read_range, open, handles). The locator composes an in-memory Okapi BM25 index with per-language Snowball analyzers and a hardened exact-search leg. Every performance-relevant default traces to a registered, mechanically adjudicated study; one default has already been retired by its frozen prediction.","date":"2026-08-22","doi":"10.5281/zenodo.22061819","doiUrl":"https://doi.org/10.5281/zenodo.22061819","repository":"https://github.com/mentu-ai/mentu-navigator","supplementTo":["10.5281/zenodo.21960138","10.5281/zenodo.21969901"],"license":"Open source"}]