Two hours, hands-on. You start with an AI agent that anyone can make do anything, and you leave with one that can't exceed the permissions of whoever is signed in.
Here's the first thing you'll do. It issues a refund against a support API with no login, no token, and no identity of any kind:
$ ./bin/lab verify 1
✓ the refund reached your API and succeeded, with NO token at all
No login, no identity, no session. That is the problem you are here to fix.
It works because in lab 1 nothing is checking. Over the next five labs you close that, one control at a time, and you type all of it yourself.
This is the hands-on companion to tyk-mcp-token-exchange, which is the same system as a 30-minute demo. That repo shows you a finished thing. This one has you build it.
An AI support copilot that calls APIs on behalf of a signed-in person, where the gateway swaps that person's broad login token for a narrow one, minted per tool call, scoped to a single action, good for thirty seconds.
By the end, the same refund request succeeds for one user and is refused for another, with no code change between them. The only difference is one attribute on one user.
flowchart LR
REQ["a rep asks<br/>the copilot"] --> G1
G1["<b>GATE 1</b><br/>may this <i>person</i><br/>use this tool?<br/>reads: entitlements<br/><i>lab 3</i>"] --> EX["token exchange<br/>narrow it<br/><i>lab 4</i>"]
EX --> G2["<b>GATE 2</b><br/>may this <i>token</i><br/>do this operation?<br/>reads: scope<br/><i>lab 5</i>"]
G2 --> UP["your API"]
G1 -->|no| D1["403<br/>no token minted"]
G2 -->|no| D2["403<br/>token insufficient"]
Two gates, two questions that sound alike, two different claims. Learning why they're different is most of the two hours.
Please do the prep the day before. It's about fifteen minutes: install four tools, get a free trial licence, and pull the images at home. Thirty laptops pulling 2 GB on conference wifi is a bad afternoon for everybody.
./bin/lab checkGreen across tools, licence and cluster means you're ready.
| Lab | You add | Time | |
|---|---|---|---|
| 0 | Setup | a running cluster | 15 min |
| 1 | The problem | nothing, on purpose | 15 min |
| 2 | Identity arrives | authentication | 20 min |
| 3 | Gate 1 | per-tool entitlement checks | 25 min |
| 4 | The exchange | token narrowing | 25 min |
| 5 | Gate 2 | per-operation scope checks | 15 min |
| 6 | Who did what | the audit trail | 5 min |
Lab 2 is the one that surprises people. You turn authentication on, and alice can still issue refunds. Knowing who is calling and limiting what they may do are separate problems, and most agent deployments have solved exactly one of them.
You will, and it's planned for. One command puts you wherever the room is:
./bin/lab goto 3 # jump to the start of lab 4
./bin/lab status # which controls are live right now?
./bin/lab verify 3 # did lab 3 actually work?
./bin/lab users # what does each user's token claim?goto overwrites your data/ with a known-good configuration and applies it. Nothing you've typed is precious. The learning is in the next lab, not the last one.
Run ./bin/lab users at any point. It's the shortest explanation of the whole design:
alice entitlements='customers:read' scope='openid customers:all'
bob entitlements='customers:read refunds:write' scope='openid customers:all'
Same scope, different entitlements, one token each. Which claim you believe is the entire workshop.
The labs are written to be handed to a room. If you're facilitating, docs/facilitating.md has timings that survived contact, the three things that always go wrong, and what to say at each transition so the point doesn't get missed.
It's also fine to work through this alone at your desk. The verify steps tell you whether each control actually took effect, so you get the same feedback a facilitator would give you.
├── labs/ the six labs, in order
├── bin/lab check · status · goto · verify · users
├── stages/ known-good config at each stage, for catching up
├── solutions/ the finished configuration
├── data/ your working copy — this is what you edit
├── docs/ prep, facilitating, stretch goals, limitations
├── k8s/ manifests and the Tyk custom resources
├── services/ the copilot and the MCP tool server (Go)
└── scripts/ numbered, run in order
stages/ is generated from solutions/ by ./bin/regen-stages.py, so the catch-up path and the finished config can't drift apart.
Say these out loud rather than letting someone find them. Full list in docs/limitations.md.
- It does not stop prompt injection. Nothing here prevents an agent being tricked into calling a tool. It ensures that when it does, the call carries only the permissions of the person it's acting for. That's containment rather than prevention, and it's the most common misreading of the whole design.
- This is impersonation, not delegation. The exchanged token keeps the rep's
suband records the gateway asazp. RFC 8693's formal actor chain viaactis not minted by Keycloak, Okta or Auth0. Ping and Curity do. - The identity provider is on the critical path. Keycloak goes down, tool calls fail. That's the honest trade for keeping zero long-lived credentials anywhere.
Apache-2.0. Run it, fork it, teach from it, rip the labs apart and rebuild them for your own stack.
Credentials in this repository (Acme-Demo-2026!, Workshop-2026!, acme-demo-exchange-secret) are fixtures for a local kind cluster, published on purpose so the labs run with no setup. They're safe to read and useless anywhere else. Your Tyk trial licence is the one thing you supply yourself, in .env, which is gitignored.
data/realm-acme.json also contains an RSA private key (acme-static-rsa), and that is deliberate. Keycloak re-imports the realm into an empty database on every restart, which would normally mint fresh signing keys and leave the gateway holding a stale JWKS — every user failing with no matching KID found in any JWKs. Pinning the key keeps the kid stable across the restarts labs 3 and 6 ask you to do. It signs tokens for a throwaway local Keycloak and protects nothing. Generate your own before this realm goes anywhere real.
./scripts/99-teardown.sh # remove the workshop, keep the cluster
./scripts/99-teardown.sh --cluster # delete everything