"We delivered every feature on the list, and we billed exactly the man-months in the contract."
Both statements can be true while the project was a commercial failure. That is the structural problem with man-month contracting: it pays for effort spent and features handed over, not for anything the business can actually bank.
The consequences compound in three directions.
Incentives point the wrong way. Under hourly or man-month billing, a vendor who works faster earns less. A vendor who scopes a feature more elaborately earns more. Every efficiency gain — including the ones AI tooling now makes routine — is a revenue cut for the vendor who achieves it. No one designs a contract to punish efficiency, but that is what this one does.
The buyer holds all the delivery risk. A vendor can ship one hundred percent of an agreed specification and still leave the client with software no one uses. Under a man-month contract, the invoice is valid either way.
Agencies and SIers get commoditized. If what you sell is developer hours, the only lever a buyer has is the hourly rate, and the negotiation runs in one direction.
The alternative is not "trust us more." It is to write measurable outcomes into the contract itself. That is harder than it sounds, and most attempts fail for the same reasons. This article covers why they fail, then a four-tier model that we use with Japanese clients, and the four elements every outcome clause needs to survive contact with reality.
Why most outcome clauses fail
Before the model, the failure modes — because an outcome clause that cannot be adjudicated is worse than no clause at all. It creates a dispute instead of a project.
- No baseline. "Reduce processing time by 50%" means nothing if no one measured processing time before the project began. Baselines must be captured and signed off before development starts, not reconstructed afterwards from memory.
- No named measurer. If the metric lives in the client's analytics stack and the vendor has no access, the vendor cannot manage what it is being held to. If the vendor self-reports, the client cannot trust the number. Both sides need to agree who reads the instrument and when.
- No attribution boundary. Conversion rate can fall because a competitor cut prices, because the client changed the pricing page, or because the software is slow. A clause that does not separate what the vendor controls from what it does not is unenforceable.
- No consequence. A metric written into a contract with no linked outcome — acceptance, payment tranche, remediation obligation — is a KPI in a document, not a contractual term.
The 4-tier outcome model
The tiers run from what the vendor controls almost entirely to what the vendor influences but shares with the business. That gradient matters: it determines how tightly each tier can be bound to money.
- Tier 1 — Delivery velocity. Deployment frequency; lead time from agreed requirement to running software in a client-accessible environment. This is the tier a vendor controls most directly, which makes it the safest to write as a hard commitment. At VAON, the working target is a 40% reduction in the time from requirement to working software, achieved by putting AI tooling — including Anthropic's Claude Code — inside the specification-to-demo loop rather than only inside the coding step. The commercial value is not cheaper code. It is that a wrong assumption surfaces in a demo in week two instead of in UAT in month four.
- Tier 2 — System performance and reliability. API response time percentiles under a defined load profile; defect escape rate by severity class; automated test coverage on the modules that matter. Write these as measured values against a named test environment and a named load profile, or they mean nothing. A p95 latency figure with no stated concurrency is not a commitment.
- Tier 3 — Adoption and usability. Task completion time for the actual operators; active user ratio against the intended user population; drop-off at the steps you care about. This tier requires the vendor to stay past handover, because the instrumentation only produces signal after real use. It also requires the client to grant access to that instrumentation.
- Tier 4 — Business impact. Conversion rate, support ticket volume, hours removed from an operational process, total cost of ownership. This is the tier everyone wants in the contract and the tier that most often produces disputes, because the vendor shares control with pricing, marketing, staffing, and the market itself.
Our recommendation on Tier 4 is deliberately conservative: bind it to acceptance criteria and to the design of the discovery phase, not to a payment penalty, unless the vendor genuinely controls the whole loop. Where the mechanism is fully in scope — for example, automating a defined class of support enquiry, or removing a measured number of hours from a specific workflow — Tier 4 can carry real commercial weight. Where it is not, forcing it into the contract produces a clause both parties will argue about instead of a project both parties will finish.
The four elements every outcome clause needs
For each metric you write into a contract, write these four alongside it. If any one is missing, the clause will not hold.
- Baseline. The measured starting value, the date it was measured, the method, and a signature from both sides. Capture it during the specification phase.
- Measurement method and owner. The instrument, the query or event definition, the environment, and the named party who runs it. If it is the client's tool, the contract grants the vendor read access.
- Measurement window. The date the measurement is taken and over what period. "Thirty days of production usage beginning at Go-Live plus two weeks" is a clause. "After launch" is not.
- Consequence and exclusions. What happens when the number is met and when it is missed — acceptance, a payment tranche, a remediation obligation with a fixed scope — and the named events that suspend the clause, such as the client changing the underlying business process or a third party altering an API.
Where the commitment actually sits
Outcome metrics do not fix the largest single source of loss in custom software, which is not slow coding. It is a specification that was wrong or incomplete, discovered late. Roughly half of failed projects trace back there.
That is why our own contractual commitment sits upstream of the tiers rather than only downstream of them. VAON runs a fixed-price specification phase — requirements definition, screen design, database design, architecture, and a fixed quotation for the build — with a defined duration and a published price. If rework then occurs during the build and its root cause traces to an error or omission in the specification VAON wrote, VAON bears that cost.
The boundary is written explicitly, because a commitment without a boundary is not a commitment, it is an unpriced liability. Errors and omissions in our specification are ours. Changes to the requirement after sign-off, changes in underlying conditions such as a third-party API or a change in law, and business rules that were not disclosed when we asked for them in a recorded session, are the client's. Every rework ticket is classified against those four categories with a written record, and every out-of-scope request goes through written change control: quotation and schedule impact in writing, client approval in writing, before work starts. You will never receive an invoice for work you did not approve.
That is the same discipline the four tiers depend on. Metrics can only be adjudicated if the specification they refer to is unambiguous and if there is a paper trail behind every decision.
Security and data handling
Outcome commitments assume the system stays available and the data stays where it is supposed to. We deploy into the region the client requires, including Tokyo, when data residency is a condition. Our own product OneBot — a production RAG-based assistant with LINE integration — was built with APPI-conscious data handling as an architectural constraint, so this is a problem we have solved in our own product rather than only in a policy document. Severity classes and response commitments, including a one-hour response target for critical incidents, are defined in the contract rather than left to goodwill.
From buying man-months to buying outcomes
The purpose of a software project is not to consume a budget or to tick items on a specification sheet. It is to remove an operational bottleneck and produce a measurable commercial change.
If you want to structure your next contract around outcomes, we will map the metrics with you, tier by tier, before anyone quotes a price.
🤝 Consult on outcome-driven contracting and OEM partnerships — Contact VAON to design the metric set for your project, receive an integrated quotation, or join our B2B OEM / technical partnership programme for agencies and SIers.