Identity
Security
Posture
A one-week hackathon prototype, now shipping in Microsoft Entra. Three versions, two rebuilds, and every design decision here traces back to the research, customer feedback or review that drove it.
Not everyone needs the same keys
A company hands out keys to everyone who works there — staff, executives, contractors, guests, and the automated accounts that run jobs with nobody watching. Each group needs a different set. Give anyone more than they need, and you widen what an attacker inherits if that account is taken over.
Admins, frontline staff, executives, engineers, guests, automated accounts. The rules that are right for one group are wrong for another — and every company groups them differently.
For each group, a list of requirements every identity in it must meet: strong sign-in, no permanent admin rights, a managed device. That list is the group’s good state — and the company writes it, not Microsoft.
People join, and access granted “just for today” never gets removed. The agent checks every identity against its own group’s list, spots the ones that slipped, and handles the safe fixes itself — asking first for anything risky.
“Identity security posture” is the answer to one question: how many identities currently meet the bar their own group is held to — and what is the most valuable thing to fix next?
Everyone built the missing parts themselves
There was no single place for any of this. People pieced the picture together one tool at a time, with no shared definition of “healthy” to measure against. The closest thing was one security score for the whole company — not a good state per kind of account. So the teams who took it seriously built the missing parts by hand.
Nothing sorted accounts for you. An admin hand-picked users one at a time from a spreadsheet-like table and built every group themselves.
No built-in way to tell someone before a change hit them — so teams hand-built their own emails, timing and opt-outs.
Every rollout and every rollback was manual. If something broke at 3 a.m., a person had to undo it by hand.
Every company was solving the same problem privately, by hand — and only the teams with time to spare kept up with it. Even then, the two hardest questions were still unanswered: which change matters most, and is it safe to make?
A hackathon, and an unusual ask
A product manager on the Conditional Access team went looking for someone who could do two things at once — design the experience and build it in code. In one week, alongside my day job, that became a working prototype.
> Demonstrate the end-to-end Good State framework — from defining the good state, running a campaign, drift detection and automated remediation.
Range, from day one — one person carrying a rough idea to a working build in a week, with no handoff in between.
What one week actually looked like
Built beside my day job. Rough, but the entire loop already worked end to end — analyse, group, measure, plan the fix.
One button. Nothing existed until the agent had read the company’s set-up.
It sorted every account into suggested groups, and flagged the ones at risk.
Then one group against its own checklist, requirement by requirement.
And a phased rollout, drafted for you — pilot first, everyone last.
An AI could read a company’s set-up, sort its accounts into groups, and name the riskiest one. That was the whole bet — and it held.
Tiles, a donut and a wall of text, buried inside another product. It could tell you something was wrong. It couldn’t help you do anything about it.
“Zava” is a fictional demo company — every figure here is placeholder data. Look at the design, not the numbers.
Where every decision is written down
One private knowledge base holding every PRD, note, transcript, customer session and spec — the receipt for everything else in this deck.
Nothing lives in someone’s head or a lost thread. Every input lands in one searchable place, wired together, so the agents and I always draw from the same source.
A paper trail. Every change traces back to a note — the decision, who drove it, and the date.
This is the difference between “I think we changed it because…” and opening the note and reading the date.
Leadership made a bet on it
The build drew real attention. Leadership asked us to keep going and turn it into something we could put in front of real customers — a one-week prototype became a staffed product on a path to preview.
> Identity Posture Agent: productizes the customer’s homework — risk analysis, phased rollouts, assessment, planning, remediation — so the org doesn’t have to do it manually.
Five things real customers asked us to build
Once real companies had it in their hands, they told us what to build next. Each of these was designed because someone asked for it by name — and every card below opens the actual conversation it came from.
Notice what they have in common: not one is a new capability for the AI. Every one is about giving the human a way to stay in control of it.
The account you must never touch
Every company keeps one or two emergency admin accounts locked away, for the day everything else fails. If an automated agent ever tightens security on those, nobody can get back in. So the product has to find them, and then leave them completely alone.
The agent scans, ranks by evidence and pre-selects three — then a human has to confirm before any campaign can run.
“In an incident, ambiguity around break glass is unacceptable.” — Pim Jacobs, InSpark
04-20-2026-Inspark-Break-Glass-Sync.md“Some see having a break-glass as a risk… Sometimes to get a certification, they don’t want to have a break-glass.” — Adnan Hendricks. It could not be one rule for every country.
03-24-2026-MVP-Summit.mdNVIDIA uses a Privileged Role Administrator, not a Global Admin — and the first build silently refused it, with no error message at all.
private-preview-summary.md“Nice to see also the category ‘Break glass’ — which we asked on the MVP summit.” — Dinant Paardenkooper, IT-Impressive
04-15-2026-Private-Preview-Kickoff.mdThen it happened for real. At NACHA, break-glass was configured perfectly and still failed — because under pressure the admin did not know how to use it. That one incident produced ten repair items.
Teach it your company’s own rules
Generic best practice is easy to ignore, because it does not know anything about you. Customers wanted to hand the agent their own written policies so the advice comes back in their language, against their rules.
Upload a policy, then read back what the agent understood — the part customers insisted on.
“Could we point Entra AI agents at our 100+ infosec policies and highlight where documented policies don’t align with configuration?” — Kevin Watford, Accenture
03-18-2026-Accenture.mdIntesa Sanpaolo offered their whole policy bible — but wanted to verify what the model understood. Iterative correction, not blind one-way ingestion.
03-12-2026-Intesa.md“People don’t maintain documents.” — Marius Solbakken. So uploading a file could never be the only way in.
03-24-2026-MVP-Summit.mdAnd the uncomfortable one, from our own team: “no one’s using it… it’s quite a heavy lift.”
2026-04-10-AI-Driven-Good-State-Brainstorm.mdThat last note changed the brief. The same feature already existed elsewhere and had reached 13 tenants out of 2,947. So the problem was never the capability — it was earning enough trust that someone hands over their policies.
Not everyone in the company is the same
The first version shipped with one Microsoft-written definition of “admin”. Real companies looked at it and said their organisation does not divide up that way — and that they should be the ones drawing the lines.
Name your own segments, save them as drafts, and swap the whole setup in one move.
“Defining the persona in itself is often an impossible task… ‘employee’ — that’s a really poor persona.” — Marius Solbakken
03-24-2026-MVP-Summit.mdCanadian Tire already models its people across five different systems, and offered to hand over their own persona spreadsheet to seed it.
06-03-2026-Canadian-Tire.md“Hard-coded definitions are a non-starter for public preview. Must be dynamic and eventually customer-customizable.” — Aakarsh Nair, Microsoft VP
2026-04-08-Good-State-Demo-Session.mdIt became the #1 recurring request of the whole preview — raised across 20+ meetings, and still growing at the end.
cumulative-private-preview-report.mdBeing straight about this one: the screen on the left is designed and prototyped, not shipped. Only the Microsoft-defined admin segment made the first release. Everything else here is the roadmap that the evidence bought.
Being secure isn’t enough — you have to prove it
Security teams do not only need to be safe. Once a year they have to sit in front of an auditor, an insurer or a regulator and show it. So every recommendation had to be able to say which published standard it satisfies.
The standards the agent measures you against, stated on the record.
“They have to have, let’s say, 80 to 90% good score” — for the auditors. — Adnan Hendricks, on what actually drives the number.
03-24-2026-MVP-Summit.mdSanford Weinberg asked for NIST 800-171, CMMC 2.0 and insurance-regulation alignment — driven by findings that kept coming back in pen tests.
private-preview-summary.md“It adds to the conversation to leadership… we’re doing Microsoft best practice and what industry standard does that actually align to?” — Josh Kolka, Pivotal
06-01-2026-Pivotal.mdThe same customer set the design limit: rolling frameworks into the score “would completely mess with the weights… the score would be meaningless.”
06-01-2026-Pivotal.mdSo frameworks label the work — they never get blended into one number. That one sentence from a customer settled a decision the team had been arguing about internally. Worth flagging: of the 22 frameworks in the design, about six were named by real customers; the rest we inferred.
“Every software should always have a rollback button”
This is the simplest thing anyone asked for, and the hardest to argue with. Before a customer will let software change who can access what, they want to see what it will do, stop it halfway, and put it back.
There is no screen to show you yet. The demand is proven and the constraints are agreed — the interaction itself is still being written.
- Settled: what has to be reversible, and who is allowed to reverse it
- Open: whether “undo” means one click back, or one phase at a time
- Open: what happens to access that changed after the campaign ran
- Shipping around it meanwhile: excluded accounts up front, and a rollout that pauses itself if sign-ins start failing
“Every software should always have a rollback button.” — customer session, June
06-12-2026-Ava.md“Without knowing exactly what happens to the end user, I’m not going to run the campaign.” — Preetam Deshpande, NVIDIA
04-30-2026-NVIDIA.md“Our VIPs and executives must never be accidentally included early. Exclusions must be explicit.” — Intesa Sanpaolo
03-12-2026-Intesa.mdBanco do Brasil and NACHA had already built their own safety nets by hand — creating spare admin accounts before any change, in case they locked themselves out.
03-10-2026-Banco-do-Brasil.mdAnd the honest status: our own engineering lead wrote it down — “the rollback button is in the mocks but there’s no implementation.” The guardrails around it are real. Rollback itself is the one thing on this list I cannot show you working.
Rebuilt properly — and it fell flat
A page of its own, built to the house style: four numbers and a table.
Out of the agent and onto a surface of its own. Consistent with everything around it, legible, and technically correct.
It looked like every other settings page in the product. Nothing on it suggested an AI had actually reasoned about your company.
Internal feedback. External feedback. One answer.
Over two weeks we got feedback on version 1 from two places — inside the company and outside it, and neither group knew what the other had said. Grey text is my plain-English translation.
“Good State needs to be more ‘AI-driven’ — current Phase 1 is deterministic logic that doesn’t showcase AI value.”Good State was the product’s internal name. The complaint: version 1 followed fixed if-this-then-that rules, so nothing it did actually needed AI.
2026-04-10-AI-Driven-Good-State-Brainstorm.mdThe decision out of that room: Phase 2 must include at least one inarguable use of AI — not AI as a label on rules that already worked.The next version had to do at least one thing that genuinely could not be done without AI.
same note · Decisions tableThe bar that came with it: if you could answer it with a database query, it did not count as AI.If a plain search of data you already had could produce the answer, it was not AI.
same note · “where does AI actually add value?”“Need human approval > need to follow change control.” — Pivotal. Auto-starting anything was a non-starter in every regulated company.Banks, hospitals and government suppliers must approve and record every change they make. Software that starts on its own breaks that.
private-preview-summary.md · 14 of 24 sessionsCustomers could not find how to start, and did not realise “Analyze my tenant” had to run first. — NVIDIA, EvidiThe one button you had to press first was not obvious, so people landed on an empty screen.
private-preview-summary.md · pain points“I don’t like the wording of ‘cohorts’. I would rather call them personas.” — Evidi, InSparkWe invented a new word for something customers already had a word for.
04-15-2026-Private-Preview-Kickoff.mdThey look impossible together: leadership wanted more AI, customers wanted less automation. But each is describing a different half of the same job, and one move satisfies both: let the AI do the thinking, and leave the deciding to the human. That sentence is version 2.
Where it came from, and where it landed
A tab inside another product. Tiles and a donut.
Its own page — but four numbers and a table.
A headline, a trend, and three ranked moves.
The dashboard became a sentence — “your privileged tier is the biggest opportunity to reduce risk”, i.e. start with the people who can open the most.
A score you can watch move — one trend line, so progress between runs is visible at a glance.
The table became a ranked list — three moves, best first, instead of a dense table to read.
The company’s own context on screen — its industry, its rules, its size, so the advice reads as specific.
From a flat list to one ranked move
One headline read, then three moves in priority order — not an undifferentiated list.
Prioritisation is the product. A ranking of the few moves that close the most exposure beats a longer, more complete backlog.
Collapsed the flat list into one ranked set of recommended actions, led by the single highest-impact move — with the share of risk it closes next to it.
The admin starts on the move that closes the most exposure, instead of deciding where to begin.
The AI got a bigger job. A person still presses go.
Version 1 used AI to check things. Version 2 hands it the judgement: who moves first, in what order, when to stop. That is exactly why a person now has to approve it.
It scanned and ticked four boxes. Then three phases and a Start button, with no reason for that order.
It plans the waves, argues for them in plain language, and monitors them. Every wave still waits for a click.
It picks the order. Wave 1 is “the admins with the lowest predicted disruption” — the agent’s call.
It predicts the fallout before anyone moves. A forecast, not a database lookup — the exact bar leadership set.
It drafts the comms and sets the date. It writes the message admins receive and proposes a deadline with a runway.
It stops itself. It watches whether people are still getting in, and pauses on its own if that dips.
Show them exactly who this hits
The preview an admin sees before anything runs.
Admins won’t approve a change to powerful accounts if they can’t see how far it reaches. So we showed them — before anything runs.
- Who’s affected — three people in this wave, none of them new to the process.
- The exact message each person will receive, in the words they’ll actually read.
- The safeguards, written out — emergency accounts excluded, each phase starts only on a click, pause and restore at any time.
Automation earns trust by showing what it’s about to do — and by knowing what not to do on its own.
The customer who broke our assumption
Every organisation keeps a couple of emergency accounts — the spare keys that get you back in when everything else locks you out. Our rollout had to leave them strictly alone. We assumed they always looked a certain way. A preview customer showed us they don’t.
Emergency accounts, found automatically and excluded by default — with a human confirming.
A preview customer’s emergency accounts didn’t match our rigid assumption. Our rollout would have either missed them — or locked the customer out of their own company.
The product now finds emergency accounts itself, including ones defined by group, shows them to the admin to confirm, and excludes them by default.
The most valuable feedback we got wasn’t about the interface. It was about an assumption underneath it.
Every change, and who drove it
Validated in the open
Through Private Preview the direction was tested across 20+ engagements — enterprises in finance, energy, tech, government and education, plus partners, MVPs and analysts. Three themes shaped nearly every decision.
- Customisation is non-negotiable — one company’s definition of “good” is never another’s.
- Trust is earned in steps — admins climb a ladder. Few press “accept” at first, but they act on what it tells them.
- Posture is cross-functional — it can’t live in one team’s silo.
What actually shipped
code
The screenshots show a fictional demo company built for prototyping. Every figure inside them is placeholder data used to exercise the design — none of it is a result, and I make no claims from it. The four above are the real ones.
PRD to production, all in code
There was no handoff. The prototype was real, typed code from the first mock — so engineering shipped the same components instead of rebuilding from a picture. It is the only reason a one-week hackathon build could become a shipping product.
Because the design and the build were the same artefact, the design couldn’t drift from what launched. No redlines, no rebuild, no lossy translation.
Every input — briefs, notes, transcripts, research, customer calls — lives in one linked, access-gated knowledge base the agents read and write. It is why every decision in this deck has a source.
How the work actually got made
As lead designer I sat between three PMs, engineering and the wider Entra design org — and the loop below is how a rough idea became production code.
Each owned a slice — risky users and the posture dashboard; phased rollouts and the customer-facing analysis; tenant analysis, first-run and break-glass. I designed across all three.
Embedded throughout for feasibility, not just at handoff. They took my high-fidelity prototypes and shipped them as production code.
Weekly syncs with the other Entra designers, keeping this consistent with a design language much bigger than my one product.
A designer who ships ambiguity
A one-week hackathon prototype is now shipping in Microsoft Entra. What carried it there wasn’t one good idea — it was three versions, two rebuilds, and evidence I was willing to act on.
- Range — I design and build, carrying an AI idea from zero to shipped.
- Craft — the hardest feedback I got made the product; I rebuilt rather than defended.
- Trust is a design material — previews, guardrails and a human in the loop are what earn the right to automate.
- Prioritisation is the product — the value wasn’t finding more problems, it was choosing the first one.
Designing in code closed the gap between what I intended and what shipped — they were the same thing.
An award-winning hackathon, carried all the way to Public Preview.