Skip to content

ProofOne feature. Two ways to build it.

What does one featureactually cost to build?

The comparison takes one issue and builds it twice: once the conventional way, a developer driving Claude with subagents, and once by handing it to developerz.ai. No run is published yet. When one is, both receipts land here whatever they say.

developerz.ai does not write the code. The coding runs in your own AI agent, on your own servers, under your own model keys. What we do is triage the issue, dispatch it, review the pull request and ship the release.

The protocol is registered before either arm runs. Whatever the result says is what goes on this page.

The head to head

Two receipts for the same job.

The feature

The feature is named in the protocol, before either arm runs

  • Same repository
  • Same starting commit
  • Same acceptance criteria
  • Same model

The run happens in a public repository that is not ours.

venue.repositoryNamed in the protocol, before either arm runs

Not our own monorepo, so you are not being asked to check our word against our own history. Both arms land in the same public venue and both timelines are readable by anyone.

Frozen before either arm runsRegistrationDate pending
The feature and its acceptance criteria
Named in the protocol, before either arm runs
Starting commit, the same for both arms
Named in the protocol, before either arm runs
Model, identical in both arms
Named in the protocol, before either arm runs
Control arm prompt and subagent setup
Named in the protocol, before either arm runs
Stopping rule per arm: done, and failed
Named in the protocol, before either arm runs
Public venue, not our own repository
Named in the protocol, before either arm runs
Reviewer, who reads with authorship hidden
Named in the protocol, before either arm runs

No figures below. The receipts keep their labels and their sources, so you can see what will fill in.

Conventional approach

Developer + Claude

A developer drives Claude with subagents and stays at the keyboard throughout.

Model spend

measured at the run

Model provider console. A fresh API key for this arm, used for nothing else.

Not the full cost of delivering this feature

Calendar timenot measured

Will come from GitHub timestamps: issue opened, pull request opened, pull request merged

Human interventionscounted at the run

Public PR timeline. Anyone can recount it.

Labour, subscription allocation and servers are not included

Inspect the receiptwhere the figure comes from

Where this figure comes from

Primary instrumentModel provider consoleprovider_usd

measured at the run

The provider attests this spend, so you do not have to trust our database. This is the figure that counts.

Not applicableSecondary instrumentour_rows_usd

Not applicable

This arm runs on the developer's own machine. We hold no rows for it at all, so there is no figure to be missing. Nothing of ours fails to record.

The gapgap_usd

There is one instrument for this arm and it is the provider's. Nothing is hidden and nothing is missing.

no second instrument

The work itself

Pull request

not opened yetPending

Session transcript

not recorded yetWe publish it in full

A person driving an agent does not reproduce run to run. Publishing the whole transcript is the only mitigation we have, so it sits beside the pull request, not below it.

Orchestrated approach

developerz.ai

The issue is filed and the platform takes it from there. The coding runs in your own AI agent, on your own servers, under your own keys.

Model spend

higher than arm A

Model provider console. A fresh API key for this arm, used for nothing else.

Includes our orchestration tokens. Nothing netted out.

Calendar timeshorter than arm A

Will come from GitHub timestamps: issue opened, pull request opened, pull request merged

Human interventionscounted at the run

Public PR timeline. Anyone can recount it.

00Handoff: the issue is filed, then the platform dispatches it. Nothing below is asked of a person by us.

Your servers and your model keys are billed to you, not by us

Inspect the receiptwhere the figure comes from

Where this figure comes from

Primary instrumentModel provider consoleprovider_usd

higher than arm A

The provider attests this spend, so you do not have to trust our database. This is the figure that counts.

Tokens only, no dollar valueOur own recorded figureour_rows_usd

not recorded yet

No rows exist, because the run has not happened.

The gapgap_usd

A gap needs two dollar figures. Ours may have none, so it is published as not computable rather than reconciled away or quietly filled with a zero.

not computable

The work itself

Pull request

not opened yetPending

Session transcript

not recorded yetWe publish it in full

Every action in this arm is also written to the append-only audit log, which is hash chained.

Result

No result. No conclusion.

The comparison has not run. The protocol is registered first, then both arms run, then this panel fills in whatever it says. Until then nothing on this page is a measurement.

Not measured

model spend

no figure yet

Model provider console, one fresh key per arm

Not measured

calendar time

no figure yet

GitHub timestamps, issue opened to merge

Not measured

human interventions

no figure yet

Public PR timeline, recountable by anyone

n = 0. No run has been recorded, so these three cells have no sample behind them. The design tops out at one run per arm, which would not support a significance claim either.

One feature, built two ways. It is a demonstration with receipts, never a rate. Arm B includes our own orchestration tokens: you pay for those under your own keys, so they are not netted out.

Read the method

The human verdict

One reader. Both pull requests. Names hidden.

A named reviewer reads both pull requests in one sitting, without being told which arm wrote which. It stays prose. There is no score, no rating and no bar, because we have no publishable number for review quality and inventing one would be fabrication.

Reviewer

Named in the protocol

The protocol names the reviewer before the review, and authorship stays hidden until the verdict is written.

ADeveloper + Claude

no verdict yet

Bdeveloperz.ai

no verdict yet

Which one would you merge?

This answer is allowed to be neither, or no preference.

unanswered

At your volume

Your bill does not move per feature.

It moves when you add an AI dev. So the cost of a feature is whatever your monthly subscription divides into, and the more you ship the less each one carries. Here is one worked example. The estimator on the pricing page does the arithmetic for your own numbers.

This is arithmetic on published list prices, not a benchmark result. Your model spend is billed to you by your provider and is not in it.

Open the estimator on pricing

One worked example

3 AI devs, 8 features shipped in the month.

Subscription, 3 AI devsPRIMARY_TIERS.procode review and CI included, not extra
$3,000.00
Your servers, 9 enrollable boxesCI_SLOTS.perSeatbilled by your host, not by us
your host
Your model keysbyokbilled to you by your provider
your bill

Per feature, before model spend

ladder ÷ features

The subscription divided by what you shipped that month. Your model spend is billed to you by your provider and is not in it.

$375.00

Don't take our word for it

Built so you can check.

Quality is not a number we have earned the right to publish. Here are the safeguards you can inspect today.

  • A record we cannot quietly rewrite.

    What this means

    Every action goes into an append-only, hash-chained audit log. Changes leave evidence.

    Inspect the mechanism
  • The checks still have to pass.

    What this means

    CI, review and branch protections gate merges. Orchestration is not permission to skip them.

    Inspect the mechanism
  • A finding needs a source.

    What this means

    Review is cite-or-refuse. If a finding cannot be grounded in the code, it is not posted.

    Inspect the mechanism
  • You always know it is a bot.

    What this means

    The bot discloses itself. It never pretends to be a human speaking to your customers.

    Inspect the mechanism
  • Your keys, sealed and yours.

    What this means

    Model keys are envelope-encrypted: each key is sealed by its own data key, and that by a master key we cannot export. You pay the provider directly and we never resell tokens.

    Inspect the mechanism
  • Your work stays on your servers.

    What this means

    Workers run on infrastructure you control. Connections dial out. No inbound port is needed.

    Inspect the mechanism

n = 0

The limits

One feature is not a promise.

Read the limits

One feature built two ways is an anecdote with receipts. That is worth something, and it is worth less than a rate. The published limits are on the page because you would find them out anyway.

  1. 01

    n = 1.

    One feature. A demonstration, never a rate. If the next one goes the other way, that goes on this page too.

  2. 02

    We choose the feature.

    Whoever picks the task picks the winner. We publish the selection rule before looking at candidates. That mitigates it. It does not close it.

  3. 03

    We run both arms.

    Including the control we are competing against. An independent run would be worth more than ours, and we would rather say so than imply otherwise.

  4. 04

    Calendar time is not effort.

    Our clock includes queue time you would really wait. The control's clock is one person's uninterrupted attention. That framing flatters the control, not us.

  5. 05

    One more limit is published with the protocol. Writing it down before the run is the only version of it that counts, so it is not guessed here.

The method

Written down before we look.

Not yet registeredThe protocol, as a public issueNot yet registeredPublic issue, opened before either arm runs

What counts as a run.

A run counts only if the arm opens a pull request on the public venue repository that meets the acceptance criteria frozen before either arm starts, and that pull request merges. An arm that hits its stopping rule and leaves the pull request unmerged is recorded as a failure for that arm, not dropped from this page.

What the figures include, and what they leave out.

  • Money is the provider console total for that arm's own dedicated key. It is not our own rows.
  • Arm B's money includes our own orchestration spend: triage, the scout, the review lane, the babysit driver. None of it is netted out.
  • The clock is GitHub's own, issue opened to pull request merged. Arm B's includes queue time you would really wait. Arm A's is one person's uninterrupted attention.
  • Every human touch after the pull request opens is counted, in both arms.
What the protocol has to fix in advance
  1. 01The feature and its acceptance criteria, written so a third party can run them.
  2. 02The starting commit, the same for both arms.
  3. 03The model, identical in both arms. Two arms on different models are not a comparison.
  4. 04The control arm's prompt and subagent setup, published verbatim.
  5. 05The stopping rule for each arm: what counts as done, and what counts as failed.
  6. 06The public venue, which is not our own repository.
  7. 07The named reviewer, who reads both pull requests with authorship hidden.
  8. 08Where each figure comes from: the provider console for money, GitHub timestamps for the clock, the public pull request timeline for interventions.
  9. 09That the result is published whatever it says, and that one feature produces no rate.
The result is published whatever it says.

The protocol is registered as a public issue before either arm runs. If our arm loses, or stops without delivering, that is what goes on this page. GitHub exposes issue edit history, so once the numbers arrive you can check the protocol against it yourself.

Pricing

The ladder does not change.

This page compares one feature. The pricing page prices the ladder and carries the estimator. They link to each other and stay separate.

See pricing and the estimator