AI design sprints that ship a decision
What changes when the teammate in the room is a model: faster prototypes, same kill criteria, no demo theatre.
Call it an AI Design Sprint if, and only if, you will leave with a decision and a test you can run next week. Otherwise you ran a hackathon with better autocomplete.
The noun mix-up is how theatre gets a new budget line. A vanilla design sprint (GV-style, five days, one problem, one prototype, one Friday test) is a decision machine. A hackathon is a demo factory. An AI Design Sprint is a decision machine that uses models to compress the slow parts: synthesis, optioning, first-draft artefacts, throwaway landing pages, first-pass scripts. It is not "we prompted a persona." It is not "we built a chatbot in 48 hours and clapped."
We have watched the mix-up in product teams at Atlassian, in public-sector digital shops that rhyme with GovTech Singapore, in Capital One-style design orgs, and in smaller operators who need to test an offer before they hire an agency. Infinitev, an Australia and New Zealand EV and hybrid battery specialist, is the residue we point at: AI-assisted landing pages so the team could test ideas in days, not quarters. That is the shape. Not the only shape. The shape.
Three formats, three residues
Be boring and precise. Executives will thank you.
Vanilla design sprint. Time-boxed problem framing, sketching, a prototype, and a test with real users (or real operators). Residue: evidence plus a go / iterate / kill. Atlassian Team Playbook rituals and Capital One design weeks live here even when they do not use the GV trademark.
Hackathon. Time-boxed building, usually many teams, often overnight, judged by a panel. Residue: demos and energy. Useful when you already have a product owner, an integration path, and a 30-day experiment budget. Useless as a culture programme. GovTech Singapore and bank digital academies have run both the useful and the useless versions.
AI Design Sprint. Same decision discipline as a design sprint. Models in the loop to compress research piles, generate option sets, draft throwaway artefacts (pages, scripts, letters, journey variants, data schemas), and sometimes to instrument a cheap test. Residue: a decision plus an artefact you can put in front of a customer or a steering owner next week. Infinitev-style landing pages. A concierge script. A service letter. A policy explainer. A shadow workflow.
If your "AI sprint" produces a slide titled Opportunities, you ran a workshop. If it produces a repo nobody will productionise and no owner, you ran a hackathon. If it produces a test and a named decision, you can keep the name.
AI makes the first draft cheap. It does not make the decision cheap. Book the decision.
What AI is for in the room (and what it is not)
For:
- Synthesising a pile of interviews, tickets, or complaints into themes a human still has to argue with.
- Generating ten offer angles or ten journey variants so the team argues about which two to test, not which one they fell in love with.
- Drafting throwaway landing pages, emails, and scripts at the speed of a conversation. Infinitev-style tests live here.
- First-pass technical spikes: a schema, a prompt chain, a thin UI. Atlassian-style product teams already do spikes. AI shortens the spike, not the review.
- Translating a concept into the language of risk or of a GovTech-style service standard, so the constraint is visible on day one.
Not for:
- Inventing the customer. If you did not talk to one, you have a stochastic average.
- Replacing the legal or clinical or credit owner. Capital One-style and bank-style rooms still need the human who can say no.
- Production architecture. A sprint artefact is allowed to be embarrassing. A production system is not.
- Consensus. Models are fluent. Fluency is not evidence.
GovTech Singapore's better digital work already treats standards and reuse as design inputs. Adding AI does not retire the standard. It makes it faster to draft the thing the standard will judge. Atlassian's better product work already treats playbooks as defaults, not decorations. Adding AI does not retire the play. It makes a first draft of the artefact the play needs.
A four-part shape we actually run
You can do this in two days if the problem is already framed. You can do it in four if it is not. Do not do it in six. Six is how a sprint becomes a project.
Frame (with the constraint in the room)
Write the problem, the customer, and the constraint on one page. Bring the person who will stop you: risk, legal, clinical safety, credit, or ops. Infinitev-style commercial tests still have a constraint (claims, channel, brand, data). GovTech-style services have policy. Capital One-style products have credit and conduct. Atlassian-style products have customers who will tweet.
Compress (this is the AI layer)
Dump the research you already have into a synthesis pass. Generate option sets. Do not fall in love. Vote with the constraint in the room. Pick one test, maybe two. Kill the rest in public.
Build the throwaway
Landing page. Script. Letter. Concierge desk. Clickable. Whatever is thinnest. Use AI to draft. Use humans to make it true. Infinitev's useful residue was not "we have a generator." It was "we can test an offer this week."
Decide
Last afternoon. Decision-maker present. Evidence in front of them, even if the evidence is ugly. Go, iterate, or kill. Name the owner of the next test. Name the date. Circulate a page before people stand up.
If the decision-maker sends a delegate who cannot kill, you scheduled applause.
How this differs from "we added Copilot to the sprint"
A lot of teams did that. Fine. A licence in the room is not a method. The method question is: which step got shorter, and did the decision get clearer?
Atlassian squads that already demo every two weeks will feel AI as faster first drafts and faster test data. The risk is more volume into the same review. Discipline the backlog or you will drown in plausible junk.
GovTech Singapore teams that already ship against a service standard will feel AI as faster content and faster prototypes. The risk is fluent nonsense in a government voice. The standard still applies. Put a human on the sentence that a citizen will read.
Capital One-style design orgs will feel AI as faster research synthesis. The risk is skipping the customer because the synthesis felt complete. It is not complete.
Infinitev-style operators will feel AI as a way to stop waiting on an agency cycle to learn whether an offer has a pulse. The risk is testing brand-damaging claims because generation was cheap. Write the claims rule before the page.
Don'ts for people who will try to make this a festival
Don't invite twelve teams. One problem. One room. Maybe two test artefacts. Hackathons invite twelve teams. Sprints do not.
Don't start with the model. Start with the decision you need on Friday.
Don't skip users because the model "sounded like" them. Talk to three. Operators count if you are a back-office service.
Don't productionise the sprint code. You will fall in love with a spike. Put the learning in the backlog. Put the spike in the bin unless a product owner claims it the same day.
Don't run it without a kill budget. Even a cheap landing-page test needs a channel, a week, and someone to read the results.
Run-of-show
Day 0 (ninety minutes). Confirm problem, constraint, decision-maker, and the test channel (page, concierge, email, desk). If any of the four is missing, do not book the room.
Day 1 morning. Synthesis and optioning with models in the loop. Humans argue. Constraint in the chair.
Day 1 afternoon. Build v1 of the throwaway artefact. Infinitev-style page, or a script, or a letter.
Day 2 morning. Expose it to three users or three operators. Repair. Tighten the decision page.
Day 2 afternoon. Decide. Go / iterate / kill. Owner, date, metric. Residue circulated.
Atlassian can run this inside an existing cadence. GovTech Singapore can run it inside a service assessment window. Capital One-style orgs can run it inside a design week. A specialist like Infinitev can run it instead of a quarter-long campaign cycle.
The kicker: if you cannot name the decision you will make at 3pm on day two, you do not have an AI Design Sprint. You have a busy calendar.
Where Infinitev-style tests sit relative to a bank or a government room
Infinitev needed offer-pulse. A bank (Capital One-style product, or a HK/SG desk) often needs comprehension or completion: will this explanation of a decision be understood, will this shadow onboarding reduce cycle time? A GovTech Singapore team needs a service slice a citizen can finish. An Atlassian team needs a spike that informs a backlog, not a repo that becomes a shadow product.
The artefact changes. The decision does not: go, iterate, or kill, with an owner and a date. If your AI Design Sprint in a bank produces 12 chatbot concepts and no test licence, you ran a hackathon. If your GovTech sprint produces a fluent page that fails the service standard, you ran a content mill. Put the standard or the conduct rule in the room on day one, the same way Infinitev must put claims in the room.
Cost, risk, and the "we already have Copilot" objection
You already have a licence. Good. A licence is not a sprint. The sprint is the decision discipline plus a throwaway artefact plus a channel to learn. Copilot in the IDE does not test an offer. A chat window does not replace three users.
Cost the sprint like a product spike: two days of a small room, a small test budget, no platform. Risk it like a test licence (see the lean piece): cohort cap, expiry, named stopper, data rule. Capital One-style and bank-style rooms already know how to spike. GovTech rooms already know how to assess. Atlassian rooms already know how to demo. Infinitev-style rooms need permission to be unpolished in public. Give it to them, inside the claims rule.
If finance asks why this is not a hackathon, show them the Monday artefact: a decision page, not a trophy.
A decision log that survives the lift
An AI Design Sprint that does not write the decision down will be retold as a demo by Wednesday. Fluency is how models create that risk. Write the log before people stand up. One page. Same-day circulation. Hang it on an existing heartbeat at Atlassian, on a service assessment window at a GovTech Singapore-shaped team, on a design week at a Capital One-style org, or on a commercial stand-up at Infinitev.
Use these headings. Do not add a narrative appendix unless someone asks.
- Decision required by 15 on day two. One sentence. "Do we test this offer on a throwaway page this week, iterate the claim, or kill it?" If you cannot name the decision before day one, you do not have a sprint. You have a busy calendar.
- Options on the table. Two, plus stop. Not twelve concepts. Twelve is how you avoid choosing. Models will happily generate twelve. That is a reason to be stricter, not looser.
- Evidence in front of the decision-maker. Three users or three operators. Ugly is allowed. Invented customers are not. Infinitev-style pages need a behaviour (enquiry, call, bounce), not a team vote on the headline.
- Constraint that sat in the chair. Claims, credit, conduct, clinical safety, policy, brand, data. Named human, not a department. Capital One-style and bank-style rooms still need the person who can say no. GovTech-style rooms still need the standard.
- Call. Go, iterate, or kill. Date of the next look. Owner of the test channel. Metric. Dollars, even if small.
- What we will not productionise. Sprint code, prompt chains, and throwaway pages stay throwaway unless a product owner claims them the same day. Atlassian squads already know how spikes become shadow products. Write the bin rule in the log.
A practical closing script:
We will now write the log in the room. If you cannot stay, you cannot own the call. The residue leaves with the owner, not with the facilitator.
When not to use a model
AI Design Sprints fail when the model is treated as a teammate with no job description. Give it a job. Also give it a list of jobs it does not get.
Do not use a model to invent the customer. If you did not talk to one, you have a stochastic average. Synthesis of interviews you actually ran is useful. A persona generated from a prompt is costume. Talk to three. Operators count if you are a back-office service.
Do not use a model to replace the stopper. Legal, clinical, credit, conduct, claims. Capital One-style product rooms, HK and SG bank desks, NHS-shaped clinical safety, Infinitev claims: fluency is not a signature. The human who can say no sits in the chair. The model can draft the risk note. It cannot sign it.
Do not use a model on the sentence a citizen or a customer will read without a named editor. GovTech Singapore teams already know fluent nonsense in a government voice is a service-standard failure. Banks know it as conduct. Energy retailers know it as a marketing claim. Write the editor into the run of show.
Do not use a model to skip the test channel. A generated landing page that never sees traffic is a mural with better kerning. Infinitev's useful residue was speed-to-learning, not a generator. Book the channel on day zero.
Do not use a model to manufacture consensus. Models are agreeable. Agreeable is how twelve people leave thinking they decided. The log forces a call. If the room wants "more options," they are asking to avoid the kill. Give them the stop.
Do not use a model for production architecture. A sprint artefact is allowed to be embarrassing. A production system at ANZ, DBS, or Telstra is not. Put the learning in the backlog. Put the spike in the bin.
Do not use a model when the decision is already a no. If procurement cannot buy the test, if the act forbids the cohort, if the claims rule kills the offer, write the no and go home. Generating variants of a forbidden idea is theatre with a GPU.
Atlassian playbooks, Capital One-style design weeks, GovTech standards, and Infinitev-style commercial tests all survive this list. The model compresses drafts. The decision stays expensive. That is the point.
A practical day-zero check:
If any of these is true, leave the model in the bag for that step: no real user scheduled, no stopper in the room, no test channel, no claims or conduct rule on the page, no named editor for public sentences. You can still run a sprint. You cannot honestly call it an AI Design Sprint until those seats are filled.
News and insights for innovation, digital transformation, future of work and L&D leaders.
Stay ahead of learning and development, corporate innovation and digital transformation news. Plus the future of work. For leaders in AU, NZ, HK, SG, the US, the UK and Canada.
Keep reading.
CBA's CAIO booked A$200m. The scam agents are the other half.
Ranil Boteju printed a FY26 number. James Roberts is fighting the con after it already worked.
Airservices Australia RFI'd an Enterprise AI Broker. Closes 30 September.
One governed front door in a private AWS tenancy. Multi-model. IRAP PROTECTED. Outputs are not certified aeronautical info.
Anthropic leased Stage 1 in Queensland. Inference for Claude, not training.
Western Downs Digital Park near Dalby. Premier Crisafulli, Dexus, Zerra DC and Macquarie. FIRB still pending. 16–17 September.

