Case study · Design-to-code governance

Keeping AI-generated UI on-system.

AI tools could build a product screen in minutes. Nothing stopped them inventing colours or rebuilding components that already existed. I built the step that checks every screen against the design system before it ships, whoever or whatever made it, and it became part of how the engineering org builds with AI.

My role
Owner of the design system. Designed and built the governance step, wrote the model behind it, drove adoption beyond engineering. [NEED: who you worked with, by role]
Where
A cybersecurity platform company. Product details are under NDA.
Tools
Angular, TypeScript, PrimeNG, Tailwind, Figma, Claude Code.
01 · The problem

Speed without a floor.

When the company started generating UI with AI tools, the design system had a new problem. The tools were fast, but they drifted. They invented colours, rebuilt components we already had, and made visual decisions nobody had approved. Hand-built UI had the same problem. It was just slower about it.

The design system described what good parts looked like. Nothing enforced it at the moment a screen was made.

02 · Version one

The first gate was a wall.

My first version asked one question: is this in the kit? If yes, it shipped. If no, it failed.

Version oneA new screen hits a single check: is it in the kit? Yes means it ships. No means it's blocked, so anything new fails.New screenIn the kit?ShipsBlockedeven if it's goodyesno
Version one. Anything new was a failure.

It was safe, and it was wrong. It turned the design system into a ceiling. Nobody could propose anything new without a fight, so people either gave up or went around the system, which is the opposite of what a system is for.

03 · The reframe

Permit new compositions. Prevent undeclared visual decisions.

The fix was to stop asking one question and ask three separate ones.

  1. Where did this come from? An existing part, a new arrangement of existing parts, or something brand new.
  2. Is it built correctly? If not, snap it back to the system.
  3. If it's new, how far does it reach? Does it stay local, or does it touch the foundations?
The three-question modelEvery change is asked three separate questions: where it came from, whether it's built correctly, and if it's new, how far it reaches. Whether it touches the foundations is worked out by the system, not declared by the author. New work is tracked and approved; only undeclared drift is stopped.A change to the UI1. Where did it come from?ExistingNew arrangementBrand new2. Is it built correctly?YesSnap it back3. If it's new, how far does it reach?Stays localTouches foundations** Worked out by the system from the change itself, never self-declared.New is tracked and approved. Undeclared drift is stopped.
The three-question model. Labels are illustrative.

Brand new is never a failure on its own. Undeclared drift is. And the one thing an author can't declare for themselves is whether they touched the foundations, because that's the one they'd be tempted to understate. So the system works it out from the change itself.

New is never a failure. Undeclared is.

That model became the contract the rest of the tooling follows.

04 · How it works now

Snap what fits. File what doesn't. Let a designer decide.

Every screen, generated or hand-built, gets checked against the approved parts and colours. Where it fits, it snaps back to the system and ships. Where it doesn't, it files a clear, typed gap instead of quietly inventing something. A designer makes the call, the gap becomes a real component, and that component goes back into the kit the next check uses.

The governance loopAn idea is generated by a person or an AI tool, then checked against the kit. On-system work snaps to the system and ships. Anything new files a typed gap, a designer makes the UX decision, it becomes a real component, and it goes back into the kit that future checks use.IdeaGenerateperson or AI toolCheckagainst the kitThe kitparts, coloursSnap to systemthen it shipsFile a gaptyped, visibleUX decisiona designer calls itReal componentfitsnewback into the kit
The governance loop, redrawn for this site.
05 · Adoption

Letting real users find the holes.

It worked in my hands. The real question was whether it worked for people who weren't me.

Other teams' tools started calling it, and it became a named step in the engineering org's AI development workflow. The CTO sponsored rolling it out to designers and PMs, not just engineers. I moved PMs onto a zero-setup path inside the AI design tool they already used, and treated their first real runs as research, not support tickets.

That paid off quickly. A PM running a real design found colour being used to mean "safe to act on", something nobody had planned for, and it became a formal system proposal. Leadership then moved to consolidate two design systems onto mine.

Adoption is still early, and I'd rather name it honestly. The first PMs are building with it, and every genuinely useful finding that month came from someone running a real screen.

06 · What changed

From a rule nobody enforced to a step everyone passes through.

  • A named step in the engineering org's AI development workflow.
  • Other teams' tools depend on its contract.
  • CTO-sponsored rollout to designers and PMs.
  • Two design systems moving to consolidate onto one.
  • [NEED: any rough number. Teams or tools calling it, gaps turned into components, time to get a screen on-system before and after, or PMs and designers using it.]
07 · What I'd do differently

[NEED: your answer]

[Two or three honest sentences. Staff panels always ask this, and they want your answer, not a polished one.]