Featured image

Stop Waiting for the Standard Link to heading

The last two posts covered the diagnosis. Authentication has been a solved problem for a decade and most of the web still gets it wrong. And now we are attaching agents to it without having agreed on how to tell an agent from a human, or how to express what a human authorised it to do.

That standard is coming. Too many large organisations need it for it not to arrive. My guess is 12-24 months before something credible is ratified, and another 2-3 years before it is widely deployed, which is roughly how long every previous identity standard has taken.

The question is what you do in the meantime, because “wait and see” is not a strategy when the traffic is already arriving. Here is my list.

Judge the Session, Not the Call Link to heading

Start here, because it is the highest-value change and it does not depend on any standard at all.

Right now, almost every API in production authorises each call in isolation. A request arrives, the token is validated, the scope is checked, the call succeeds. The next request knows nothing about the last one. There is no memory and no opinion about sequence.

That model cannot see the attack, because the attack is the sequence.

The takeover from the last post is worth looking at again, because this is the version your rule engine sees. Same four requests, one extra column:

09:41:55  session.start           device unrecognised       risk 20
09:41:58  session.elevate         step-up satisfied         risk 35
09:42:04  profile.contact.update  email -> attacker@...     risk 60
09:42:11  credential.reset        link to the new address   risk 90  <- challenge

Every one of those calls is individually authorised and individually unremarkable. Read down the column instead and the story is obvious: nobody legitimately changes their recovery address and then immediately resets a credential from a session that just elevated itself. That is not a suspicious call, it is a suspicious story, and the story is only visible if you keep the previous chapters.

So keep them. You do not need machine learning to start. You need a ring buffer and a few honest opinions:

// Last N operations for this session, newest last.
const recent = session.trail.slice(-8).map((e) => e.op)

const DANGEROUS = [
  {
    name: "takeover",
    // an elevation, then a change of where recovery goes, then a credential change
    pattern: ["session.elevate", "profile.contact.update", "credential.*"],
    within: 300, // seconds
    weight: 70,
  },
  {
    name: "exfiltration",
    pattern: ["export.*", "export.*", "export.*"],
    within: 60,
    weight: 40,
  },
]

function sequenceRisk(trail) {
  return DANGEROUS.filter((rule) => matchesInOrder(trail, rule.pattern, rule.within)).reduce(
    (score, rule) => score + rule.weight,
    0,
  )
}

matchesInOrder is a subsequence test with a time window, which is 30 lines and no dependencies. A handful of hand-written rules for your five most dangerous flows will catch more real attacks than another year of scope tuning, and unlike a model you can explain every decision it makes to an auditor.

This is a graph problem, and it is the same graph problem regardless of whether the actor is a human or an agent. Which is convenient, because you have to solve it either way.

Make Risk a Number That Moves Link to heading

The second change follows directly from the first, and finance worked it out decades ago.

Card schemes do not treat fraud as a per-transaction pass or fail. They score. Velocity, deviation from an established pattern, merchant category, device, geography, hour of the day, all folded into a number that rises and falls as the session progresses. When it crosses a threshold, something happens. The rest of the time nobody is interrupted, which is the entire point.

APIs should work the same way, and almost none of them do. We build binary gates because binary gates are easy to reason about and easy to test. Then we discover that the binary gate is either so tight it blocks legitimate work or so loose it authorises the takeover, and there is no setting in between, because we did not build a dial. We built a switch.

Give yourself a dial. Score the session, let unusual-but-explainable behaviour raise the score without slamming the door, and let a long run of normal behaviour bring it back down. Every number below is invented for illustration; the shape is the point, and the weights have to come from your own traffic:

const SIGNALS = {
  "device.unrecognised": +20,
  "geo.impossible_travel": +35,
  "session.elevate": +15,
  "profile.contact.update": +25,
  "credential.reset": +30,
  "agent.attestation:tpm2": -20, // hardware backed, trust it more
  "agent.attestation:software": +10,
  "agent.mandate.present": -15,
  "history.clean_30d": -25,
}

// Risk decays towards baseline while nothing interesting happens.
score = Math.max(0, score * Math.exp(-elapsed / HALF_LIFE)) + weightOf(event)

if (score + consequenceOf(op) > STEP_UP) await challenge(session)
if (score > HARD_STOP) await revokeMandate(session.mandate)

Two things to notice.

The score falls as well as rises, so a session is not condemned for one odd request at the start. Come in from an unrecognised device, then behave normally for 10 minutes, and you are back to baseline without anybody being interrupted.

And the test is not the risk score on its own. It is the score plus the consequence of what is being attempted, which produces results that look wrong until you work them through. Take a step-up threshold of 80:

read an invoice       session 60  +  consequence 10  =  70   allowed
delete a customer     session 40  +  consequence 50  =  90   challenge

The safer-looking session is the one that gets challenged. That is correct. A session at 60 doing something trivial is not worth interrupting a person over, and a session at 40 about to destroy a customer record is, because the cost of being wrong is what you are actually managing. How the session got here matters less than what it is about to do.

This matters more for agents than for humans, because of consent fatigue. Interrupt a person 40 times in an afternoon and they will stop reading the prompts, then find the setting that disables them. Every approval you spend on something trivial is an approval you will not have available when it counts. Risk scoring is what lets you spend them sparingly.

Treat Agents as a First-Class Client Type Link to heading

Your identity provider currently knows about users and about applications. It probably does not have a concept for “software acting on behalf of a specific user, with a defined and bounded authority”. Add one, even if you have to model it with what you have.

That means an agent gets its own identifier, distinct from the human it acts for. It gets its own credential, revocable on its own without revoking the human. It gets its own entry in the audit log, so the record says the agent did this under this person’s authority, rather than simply recording the person and losing the distinction forever.

It also means logging and rate limiting by identity and mandate rather than by IP address. IP-based controls were always a proxy for identity, and a poor one. With agents they are worse than useless, because a thousand agents may share an egress address and one agent may rotate through a thousand.

Do this even though the standard does not exist yet. You will be modelling the same concepts the standard eventually formalises: a principal, a delegation, a constraint set, an expiry. When the specification lands you will be mapping your fields onto theirs rather than discovering you have nowhere to put the data.

Give Bots a Front Door Link to heading

If the only way an agent can use your service is to imitate a browser and click through your interface, then you have guaranteed that every good bot looks exactly like every bad one. You have destroyed your own signal and then complained that you cannot tell them apart.

Give them somewhere to go. An API. An MCP endpoint. A documented, machine-readable path through the transactions you actually want to happen. Publish what an agent is allowed to do, what it must declare, and what will get it rate limited. Make declaring yourself the path of least resistance, because right now the path of least resistance is lying, and you built it.

None of this is a new argument for me. I wrote it up in February in Building for an Audience of Machines: we have always segregated audiences on the internet using domains, subdomains and paths, and agents are simply the newest audience to turn up without an entrance built for them. The conclusion then was that maintaining REST for humans and MCP for agents is two contracts to drift apart, and that you are better off with one.

So I specified it. OpenCALL, the Open Command And Lifecycle Layer, is an operation-based, transport-agnostic protocol: one endpoint at POST /call, one envelope, and a self-describing operation registry at /.well-known/ops. An agent does not think in PUT /devices/garage-door, it thinks open the garage door, so the operation carries the intent and the envelope carries the context, correlation and idempotency key. Humans and agents are equal consumers of the same contract. OpenCALL In Practice covers what that looks like pointed at home automation, where the consumers are me, an agent in Slack, and a pile of sensors. What matters here is how much of the bot problem the same contract dissolves on contact.

The registry is the sign on the front door, and it is just a document:

// GET /.well-known/ops  (trimmed)
{
  "callVersion": "2026-02-10",
  "schemaHash": "sha256:...",
  "endpoints": ["rpc", "path"],
  "errorsUrl": "/.well-known/errors",
  "operations": [
    {
      "op": "article.read:v1",
      "executionModel": "sync",
      "sideEffecting": false,
      "argsSchema": {
        "type": "object",
        "properties": { "slug": { "type": "string" } },
        "required": ["slug"]
      },
      "authScopes": [],
      "cache": { "enabled": true, "ttl": 300, "scope": "public" }
    },
    {
      "op": "archive.search:v1",
      "executionModel": "sync",
      "sideEffecting": false,
      "argsSchema": {
        "type": "object",
        "properties": { "q": { "type": "string" }, "limit": { "type": "integer" } },
        "required": ["q"]
      },
      "authScopes": ["archive:read"]
    }
  ]
}

And the call itself is one shape, whoever is calling:

POST /call HTTP/1.1
Content-Type: application/json
Accept: text/markdown

{
  "op": "article.read:v1",
  "args": { "slug": "let-the-good-bots-in" },
  "ctx": { "requestId": "0f9c1a2e-77b3-4a1e-9c11-2b7d3f5a8e10" },
  "auth": {
    "iss": "https://idp.example.com",
    "sub": "dan@example.com",
    "credentialType": "delegated-agent",
    "credential": "eyJhbGciOiJFUzI1NiIsInR5cCI6IkpXVCJ9..."
  }
}

Because a published operation registry is a front door with a sign on it. The agent reads /.well-known/ops, discovers exactly which operations exist and what they need, and asks for one by name. It does not have to scrape your interface to work out what is possible, so it does not have to look like a scraper. You get a declared intent on every single call instead of an inference drawn from a URL pattern, which makes the sequence analysis from earlier in this post almost trivial, and it gives you one obvious place to hang a price and a required trust level per operation. Hold that thought.

The pattern to copy is robots.txt, which worked for 25 years for exactly one reason: it was easier to comply with than to circumvent, and compliance got you a better outcome. That is the only bot policy that has ever worked at scale. Every other approach has been an arms race, and the arms race is now being run against opponents who write their own code.

Be honest about the incentives, too. If your terms say agents are forbidden but your competitor’s site works fine with an agent, the customer does not abandon the agent. They abandon you. Blocking does not stop the transaction, it relocates it.

Sell the Response, Not the Page View Link to heading

Here is my genuine bet on where this goes. It is the part of this post most likely to be wrong, and the part I would most love to be right about, even if only in nature.

Advertising is finished as the funding model for the open web. Look at the actual flows. Companies pay Google to show advertisements on YouTube. People pay Google not to show them advertisements on YouTube. Money moving in both directions for opposite outcomes is not a healthy market, it is a market being harvested. So who is actually watching the ads? A shrinking, less valuable audience, on an internet increasingly read by software that will never see them and would not care if it did.

Meanwhile every service that has ever tried to monetise a bot has reached for the same two options: block it, or sell an enterprise API contract with a sales call attached. There is nothing in between, and “in between” is where nearly all the demand is.

HTTP 402 has said “Payment Required” since 1997 and has sat unused for its entire existence, waiting for a payment mechanism cheap enough and fast enough to make a per-request charge sensible. x402 is the current attempt to fill that gap, and I think it is structurally right.

But it is being marketed wrong, and the wrong framing will hold it back. Every explainer reaches for micropayments per page view, and that is not the use case. It might become one eventually. It is not the one that matters now.

NewsCorp does not make money half a tenth of a cent at a time and never will. They sell page views by subscription, at something like $24 a year, and they do that because it is vastly easier to manage. One customer, one recurring line, one renewal date, one relationship. Nobody wants to reconcile a billion individual events to arrive at the same revenue, and nobody needs to. Human attention is already monetised, and monetised competently. Do not go and disrupt the part that works.

What 402 actually buys you is an API response. That is the product.

And when the caller is a bot, the correct response is not a page. It is a JSON payload or a lump of markdown. No navigation, no cookie banner, no consent modal, no advertising slots, no layout, no 300KB of JavaScript to render 800 words of text. There is no person there. Bots need plain text, and every byte beyond the plain text is waste that you paid to serve and they paid to receive.

So the thing being sold is a clean, structured, machine-readable answer, priced per response, with no account, no sales call, no subscription and no advertisement. That is a genuinely new product, and it is one your existing publishing stack can produce almost for free, because the text already exists. You are just declining to wrap it in furniture nobody is looking at.

The exchange is two round trips. The caller asks, the server quotes, the caller pays, the server answers:

POST /call HTTP/1.1
Accept: text/markdown

{ "op": "article.read:v1", "args": { "slug": "let-the-good-bots-in" } }
HTTP/1.1 402 Payment Required
PAYMENT-REQUIRED: eyJzY2hlbWUiOiJleGFjdCIsIm5ldHdvcmsiOiIuLi4iLCJhbW91bnQ...
POST /call HTTP/1.1
Accept: text/markdown
PAYMENT-SIGNATURE: eyJzaWduYXR1cmUiOiIweGE5ZjMuLi4iLCJwYXlsb2FkIjp7Li4u

{ "op": "article.read:v1", "args": { "slug": "let-the-good-bots-in" } }
HTTP/1.1 200 OK
Content-Type: text/markdown
PAYMENT-RESPONSE: eyJzZXR0bGVkIjp0cnVlLCJ0eG4iOiIweDdmM2M...

# Let the Good Bots In

The last two posts covered the diagnosis...

Those headers carry base64-encoded JSON in both directions, so the payment requirements, the signed authorisation and the settlement receipt all travel in-band. Read the x402 documentation for the exact object schemas rather than trusting my abbreviations above. What matters for this argument is the shape: no redirect, no checkout page, no account, no human. A status code, a quote, a signature, an answer.

Risk Goes Surprisingly Well with Reward Link to heading

Now put the two halves of this post together, because this is where it gets interesting.

Once you have an identity for the caller and a mechanism to charge them, you can price. Not one price for everyone who is not blocked, but a tariff that reflects what you know about the caller and what they are doing. Risk tiers on one side, membership rewards on the other, which is a sentence I did not expect to write about bot traffic.

And this is the thought I asked you to hold. If your operations are declared in a registry, that is where the price list belongs, sitting next to the schema. Each operation carries what it costs and what the caller has to present to get it, so an agent knows the price before it commits rather than discovering it in a 402 it did not expect. A menu, in other words, which is how every other business on earth sells things.

The good bots get the good rate. An agent that identifies itself, carries a mandate, stays inside the documented paths and has a clean history is a customer, and you treat it like one. Cheaper per response, higher limits, first access, the same way every loyalty programme on the planet has worked for 50 years. Reward the behaviour you want.

The extraction case pays a rising tariff. An unidentified caller pulling your entire archive is not doing the thing you built the endpoint for, and the price should notice.

CallerWhat it presentedPer responseNotes
MemberMandate, hardware attestation, 90 days clean$0.0005Highest limits, first access
DeclaredAgent identity, no attestation$0.001The default good citizen
AnonymousNothing but a user agent$0.002Works, costs more
SweepingAnything, past the volume ceiling$0.004 and climbingNot blocked, just priced

Nobody is blocked. The first 1,000 responses cost almost nothing. Taking everything costs real money, which is exactly right, because taking everything is worth real money to whoever is taking it. And the tariff is published in the same registry as the operations, so a well-built agent can decide whether the answer is worth the price before it asks:

// A registry extension, not part of any spec. This is the bit
// somebody needs to standardise.
{
  "op": "archive.search:v1",
  "price": { "member": "0.0005", "declared": "0.001", "anonymous": "0.002" },
  "currency": "AUD",
  "escalation": { "after": 1000, "period": "24h", "multiplier": 2 }
}

And notice what you have built. The risk score that decides whether to challenge a session is now the same score that decides what that session pays. One signal, two uses, and the second one is revenue rather than cost. Every fraud team in the world already runs the first half. Almost nobody has connected it to the price list.

Payment also turns out to be an excellent trust signal in its own right. A user agent string is free to forge. A settled payment is not, and it arrives attached to something you can identify, rate limit, reputation score and cut off. You will get a better answer to “is this bot acting in good faith” from a payment history than from any amount of fingerprinting.

I think this gets big, and soon. I could be wrong about the specific protocol. I do not think I am wrong about the shape.

Build for a Standard That Does Not Exist Yet Link to heading

The last one is architectural, and it is how I am approaching my own work.

Assume you will be wrong about the details. Assume the mandate format you invent this year will not be the mandate format that wins. Then build so that being wrong is survivable: keep the identity and delegation layer separable from the business logic that consumes it. Express authority as data with a defined schema rather than as conditionals scattered through your handlers. Keep the decision of “is this actor allowed to do this” in one place that you can replace.

If you do that, adopting the eventual standard is a translation exercise. If you do not, it is a rewrite.

And then advocate. I cannot make a standard happen on my own, and neither can you, but the specifications that end up mattering are shaped by the people who bother to turn up: read the drafts, comment on them, push your vendors to commit publicly to interoperability rather than to their own registry. The absence of a standard is not a neutral state. It is an open invitation to anyone who wants to build a bad bot that looks exactly like a good one, and every month we spend in the wild west makes that easier.

The Short Version Link to heading

Assume every credential your users can recite is already public. Move to possession-based authentication and stop building your own login.

Judge sessions as sequences, not calls. Score risk on a scale that moves, and spend your human interruptions where they matter.

Give agents their own identity, their own mandate, their own audit trail, and their own front door. Make declaring yourself easier than lying.

Sell the response rather than the page view, in plain text, at a price that rewards the bots you want and climbs for the ones stripping the place bare. A bot that pays is a bot you can identify, and a bot you turn away is a customer your competitor gets.

And build all of it so it can be replaced, because the standard is coming and none of us know exactly what it looks like yet.

The bots are already here. Refusing to plan for them is still a decision, and right now it’s the one being made by default in an awful lot of places.


This post concludes a series on identity and agents. Previously: Auth Is a Solved Problem on why authentication is still broken everywhere, and Who Authorised the Bot? on the gap between agents and identity. See also Building for an Audience of Machines and OpenCALL In Practice for the API contract this all sits on.