Everyone Is Already on the List Link to heading

Your credentials are in there somewhere
Go and put your email address into Have I Been Pwned. I will wait.
For almost everyone reading this, the answer is yes. Not once, either. Your email address, your phone number, your date of birth, your residential address, your mother’s maiden name, the name of your first pet, the last four digits of a card you cancelled in 2019 - all of it has leaked from somewhere, been aggregated with everything else that leaked, and is sitting in a list that costs less than a coffee to buy.
That is the actual state of play, and it has a consequence that most systems still have not internalised. Every authentication method built on the assumption that you know something nobody else knows is now broken. Not weakened. Broken. Knowledge-based authentication is dead, security questions are a joke, and if your account recovery flow asks for a date of birth and a postcode then your account recovery flow is a public API for taking over accounts.
The genie is out of the bottle and it is not going back in. So the honest starting position for anyone building anything in 2026 is this: assume the attacker already has everything a legacy system would have asked the user to prove.
We Worked Out How to Fix This a Long Time Ago Link to heading
Here is the part that should be embarrassing for the industry.
This is not an unsolved problem. It is not an open research question. The patterns that fix it have been written down, standardised, implemented, interoperability tested and given away for free for well over a decade.
OpenID Connect landed in February 2014. It sits on top of OAuth 2.0 and does the boring, unglamorous work of telling a relying party who the user is, with signed assertions, key rotation via a published JWKS endpoint, and a discovery document so the client can work most of it out on its own.
That discovery document is one command away, on any provider worth using:
curl -s https://accounts.google.com/.well-known/openid-configuration | jq '{
authorization_endpoint,
token_endpoint,
jwks_uri,
code_challenge_methods_supported,
id_token_signing_alg_values_supported
}'
{
"authorization_endpoint": "https://accounts.google.com/o/oauth2/v2/auth",
"token_endpoint": "https://oauth2.googleapis.com/token",
"jwks_uri": "https://www.googleapis.com/oauth2/v3/certs",
"code_challenge_methods_supported": ["S256"],
"id_token_signing_alg_values_supported": ["RS256"]
}
Every endpoint you need, the signing algorithms, the supported challenge methods, and a URL where the public keys live and rotate without you doing anything. Your client fetches jwks_uri, validates the signature on the id_token against the key with the matching kid, checks iss, aud, exp and nonce, and it is done. No password ever touched your infrastructure. It is not elegant in every corner, but it is well specified, widely deployed and thoroughly attacked, which is worth more than elegance.
PKCE followed in 2015 and closed the authorisation code interception hole for public clients. It is three lines of work:
// Before redirecting the user to the authorisation endpoint
const verifier = base64url(crypto.getRandomValues(new Uint8Array(32)))
const challenge = base64url(await crypto.subtle.digest("SHA-256", utf8(verifier)))
sessionStorage.setItem("pkce_verifier", verifier)
// .../authorize?response_type=code
// &code_challenge=<challenge>&code_challenge_method=S256&state=<state>...
POST /token HTTP/1.1
Content-Type: application/x-www-form-urlencoded
grant_type=authorization_code
&code=4/0AeanS0b...
&code_verifier=<the original verifier>
&redirect_uri=https://app.example.com/callback
&client_id=1234.apps.googleusercontent.com
An attacker who intercepts the authorisation code cannot redeem it, because they do not have the verifier that hashes to the challenge the server already saw. That is the whole mechanism, it costs nothing, and it has been mandatory practice for a decade.
Sender-constrained tokens via mTLS and later DPoP closed the “possession equals authority” hole for bearer tokens. The backend-for-frontend pattern, where the tokens never touch the browser at all, has been the documented answer for high-assurance applications for years. The client ends up holding this and nothing else:
HTTP/1.1 200 OK
Set-Cookie: __Host-sid=8f14e45fceea167a; Path=/; Secure; HttpOnly; SameSite=Lax
An opaque identifier. No JWT in localStorage for any script on the page to read, no refresh token in the browser, nothing an XSS payload can exfiltrate and replay. The tokens live server side, keyed by that cookie. If your threat model says a token must never be readable by client-side JavaScript, that is a solved architectural problem with published guidance and reference implementations.
WebAuthn and passkeys then did something genuinely new: they moved the secret into hardware and made it non-exportable.
await navigator.credentials.create({
publicKey: {
rp: { id: "example.com", name: "Example" },
user: { id: userHandle, name: "dan@example.com", displayName: "Dan" },
challenge: serverChallenge,
pubKeyCredParams: [{ type: "public-key", alg: -7 }], // ES256
authenticatorSelection: {
residentKey: "required",
userVerification: "required",
},
},
})
Note rp.id. The credential is bound to that origin by the browser, not by your code, and the binding cannot be talked out of the user. Phishing a passkey is not a matter of tricking someone into typing it, because there is nothing to type, and a lookalike domain will not match the rp.id so the authenticator will not even offer the credential. You cannot read it out and you cannot post it to an attacker’s server, no matter how convincing the email is.
So the toolkit is complete, and it has been complete for a while.
Free, Freedom, or a Face and a Name Link to heading
The second embarrassment is cost, because at the bottom of the market there almost is not any. Free, freedom, and an enterprise agreement, in that order.
Free is Firebase. Firebase Authentication will handle 50,000 monthly active users for no money whatsoever. Fifty thousand. That is not a trial, it is the free tier. For a very large proportion of the applications being built right now, the entire authentication bill is zero dollars, forever.
Social login costs even less than that, if such a thing is possible. Sign in with Google, Apple, Microsoft or GitHub requires no cryptographic expertise, no password storage, no reset flow, no breach exposure on your side, and no ongoing cost. The technical requirement is registering a client with the provider and handling a redirect. That is it. You are done in an afternoon, and the hardest security problem in your application has been handed to an organisation with a security team larger than your entire company.
Freedom is FOSS. If you want to run it yourself and answer to nobody, Keycloak and Ory will both do the job, and Zitadel if you want something more recent. The licence costs nothing; you pay in operations instead, which is a real cost and an honest one. None of these are research projects. They are mature, deployed in anger, and boring in the way infrastructure should be.
And if you want an enterprise agreement and someone to shout at, Ping and Okta have been doing this for a very long time and they are good at it. That is emphatically not free, and it is not meant to be. What you are buying is a contract, a roadmap, a support line with a person on the end of it, and someone else’s name on the assurance documentation when the auditor asks. For a bank or an insurer that is worth every cent, and the cheque gets signed without much argument.
So the range runs from nothing to a great deal, and every point on it is a solved problem someone else maintains. There is no serious argument that doing this properly is expensive or difficult. It is neither.
So Why Is It Still a Hot Mess? Link to heading
I have thought about this a lot, because it is genuinely strange. The vendors are large and well credentialed. The standards are open. The open source options are good. The free tier is generous. And yet the web is still full of forms that take a username and a password over a connection you have to hope is doing its job, applications that email you your existing password when you click “forgot”, and services storing PII they had no reason to collect and no ability to protect.
There are three reasons, and only one of them is technical.
Reason One: Legacy Link to heading
The first is legacy, and it is the honest one. There is an enormous amount of software still running very large corporate services where the authentication layer was designed twenty years ago and everything since has been built on top of it. Replacing it is not a sprint. It is a multi-year programme touching every integration, every batch job, every partner connection and every internal tool nobody remembers exists. It costs real money, sometimes ten million dollars of real money, and someone has to sign that cheque.
That someone is looking at a spreadsheet. And from where they are sitting, the old system still works. Nobody is complaining. Revenue is fine. There is no incident on the board report. Spending eight figures to arrive at exactly the same user-visible outcome is a very hard business case to make, right up until the morning it becomes the only thing anybody wants to talk about. The economics of prevention are terrible and always have been.
Reason Two: Familiarity Link to heading
The second is familiarity, and this one is on us.
Username and password is what everybody knows. It is what the tutorial shows. It is what the framework scaffolds. Ask any AI coding agent to add authentication to a project and watch what it reaches for by default:
CREATE TABLE users (
id UUID PRIMARY KEY,
email TEXT UNIQUE NOT NULL,
password_hash TEXT NOT NULL,
reset_token TEXT,
created_at TIMESTAMPTZ DEFAULT now()
);
Every one of those columns is a liability you volunteered for. password_hash is a cracking target and a breach headline. reset_token is an account takeover primitive if it is ever generated with a weak random source, logged, or given a generous expiry. email is the join key that lets someone correlate your breach with the other forty they already have. The version of this table that uses an identity provider has none of those columns, because it stores a subject identifier and nothing else.
It gets scaffolded anyway, because that is what the overwhelming weight of the training data does, which is another way of saying that is what we have collectively written down a million times. The agent does it because we did it, over and over, for twenty years, in public.
I will exempt myself from this one, but only just. I have started every new project on Firebase since 2020 and I have not written a users table in years. That was not principle, it was laziness pointed in a useful direction: the free tier was easier than the alternative, and easier is the only force that has ever reliably changed developer behaviour. Which is the entire point. Make the correct path the lazy path and the argument is over.
Familiar and easy beats correct, every single time, unless someone senior enough makes correct the default. That is the whole mechanism. There is no deeper mystery.
Reason Three: Nobody Owns It Link to heading
The third is organisational. Authentication is usually treated as a feature of an application rather than as infrastructure, which means it gets built by whichever team happens to be building the application, to whatever standard that team happens to hold, on whatever timeline the product manager negotiated. Nobody owns it across the organisation. The security team finds out at the penetration test, which is to say after it has shipped, at the point where the only available options are “accept the risk” and “delay the launch”. You can guess which one wins.
There is a career observation hiding in that, and it is not flattering to the industry. In December 2022 I volunteered to help with an identity migration on the strength of having some idea of how it was all supposed to work. That was the entire qualification. Three and a half years later the job title says identity, the work is architecture, and I am the subject matter expert people book time with.
I would love to tell you that was a masterclass in career strategy. It was not. It is what happens when something this important is owned by nobody: the person who puts their hand up and reads the specification becomes the expert, because the bench is that thin. If you want to be valuable in a large organisation, find the thing everyone depends on and nobody wants, and learn it properly. There is a queue of one.
Zero Trust Was Not Enough Either Link to heading
The industry’s answer to all this for the last several years has been Zero Trust, and Zero Trust is right as far as it goes. It is also routinely misquoted, usually as “stop trusting the network”, which is a fragment of it at best.
Zero Trust means trusting nothing that is not explicitly part of the server or the service. Not the network, but not the client either. Every client is untrusted: on the corporate network or off it, managed or unmanaged, yours or somebody else’s. And for a public client you have to go one step further and assume the code has been modified, because you handed it to a machine you do not control and somebody will pull it apart. The payload arriving at your endpoint may not be one you created, may not be one you expected, and may not be one your own client is even capable of producing.
That gives you three working rules. Ignore the unexpected: an unknown field is dropped, never coerced, never persisted “just in case”. Mandate the expected: a missing required claim is a rejection, not a default. Fulfil only the known contracts: an operation that is not in the published contract does not exist, and neither does a combination of parameters you never specified.
And at some point up that chain, validating the shape of a payload stops being enough and you need the payload signed. Not because the contract might be wrong, but because you need proof it came from something you would have trusted in the first place, and you need to still have that proof tomorrow when somebody asks who sent it.
Which is a risk decision with a technical control attached to it, tier by tier:
| Risk | What you require | Why |
|---|---|---|
| Low | TLS, strict schema validation, unknown fields dropped | The contract is the control |
| Medium | Plus sender-constrained tokens, short expiry, idempotency keys | Theft of a token stops being sufficient |
| High | Plus signed payloads from a hardware-held key, replay windows, retained proof | Non-repudiation, and no compromises |
High risk means high security. Not high security with an exception for the legacy client that cannot sign, because that exception is the only door anybody will ever use.
But notice the load-bearing assumption underneath all of it: that when you verify, the verification means something. Zero Trust tells you to check the credential rather than the network location. It does not help you when the credential material itself is public.
And that is where we are. Here is the chain, written out, because it looks much worse on paper than it does in a threat model workshop:
attacker buys combo list -> email + DOB + mobile + address
calls telco, quotes DOB -> SIM ported to attacker's device
requests password reset -> OTP delivered by SMS to that device
answers "first pet" from a 2019 -> security question satisfied
breach of a forum you forgot
-> account, and every account that
recovers through this one
Every single step in that chain succeeds against a system that is doing exactly what it was designed to do. Nothing is exploited. No vulnerability is used. The controls are working and the outcome is still a takeover, because each control is checking a fact that stopped being secret years ago.
If the second factor is an SMS to a number that is on a list, and the number can be moved to a new SIM with a phone call to a call centre that will verify the caller using a date of birth that is also on a list, then you have not built a second factor. You have built a longer first factor. The chain is only as strong as the weakest identity proofing step in it, and for most organisations that step is a human on a phone with a script and a queue length target.
So the rule I would write on the wall is short: stop trusting anything a user can tell you. If it can be typed, spoken, emailed or read off a screen, treat it as public, because it probably is. What is left is possession of something that cannot be copied and cannot be exported - a key in hardware, bound to an origin, that signs a challenge and never leaves the device it was born on. That is the only category of proof that survives contact with the current threat environment.
Everything else is theatre with a compliance tick next to it.
What Actually Needs to Happen Link to heading
None of this requires new invention. It requires doing the things that already exist.
Stop building your own. Whatever you are about to write, one of the four or five identity providers named above already does it, has already had it attacked by better people than either of us, and will run it for free or nearly free. There is no competitive advantage in your login form. There is only downside.
Stop collecting what you cannot protect. If you do not need a date of birth, do not ask for one. Every field you collect is a field you can leak, and the aggregate of all those leaks is precisely the problem described at the top of this post. The safest PII is the PII you never had.
Move to possession-based authentication, properly. Passkeys as the primary factor, not as an optional extra buried in account settings that three percent of users will ever find. If you are running a consumer service, the default enrolment path should be a passkey, and password should be the fallback you are actively trying to retire.
Sender-constrain your tokens. A bearer token is a bearer instrument; whoever holds it is the owner, and that is a terrible property for something that gets logged, cached, copied into a support ticket and pasted into a chat window. Bind it to a key. Make theft insufficient.
And keep tokens away from the browser when the assurance requirement is high. The pattern exists, it is documented, and it works.
The Part That Actually Worries Me Link to heading
All of the above is a decade-old argument that I should not still be making, and I would not have bothered writing it down except for one thing.
We are now, at some speed, handing this stack to software that acts on our behalf. Agents are already logging in, already clicking through consent screens, already holding credentials, already making purchases. They are doing it through interfaces designed on the assumption that there is a human at the other end who will read the warning, notice the anomaly and hesitate before approving something strange.
There is not a human at the other end anymore. And we have not agreed on how to tell.
That is the scary part. And it’s also the next post.
This post is part of a series on identity and agents. Next: Who Authorised the Bot? on what identity looks like when the user is not a person, and Let the Good Bots In on what to do about it.