Klarna AI Assistant
The most-cited enterprise AI deployment in the world, and the one with the thinnest published architecture. Klarna has never released an engineering blog, conference talk or technical paper on this system. Everything below is reconstructed from three vendor case studies, one CEO podcast, one Klarna job advert, SEC filings, and one journalist's hands-on teardown — and every box is marked confirmed or inferred accordingly.
Read this before you use these diagrams
Siemiatkowski, on the record: “this is one time in my life that I actually feel a bit cagey about telling too much about the secret sauce, because I actually think about this as a fairly important strategic advantage.” Sequoia's companion piece confirms Klarna has not disclosed how it was built “beyond using a form of RAG.” Widely-circulated “Klarna architecture” diagrams showing microservices, confidence scores and handoff payloads originate in vendor content-marketing blogs and one LLM-generated database entry. They are not sourced. This page marks the boundary explicitly.
System context
Who talks to the assistant, what it talks to, and where the hard boundaries sit. Note the two things most reconstructions get wrong: the human agent tier is a parallel track, not a fallback; and generative AI is explicitly walled off from credit underwriting.
One corpus, two consumers
The retrieval corpus is the documentation written for human agents. Siemiatkowski: “our agents have better tools today to be successful in helping the customers as does the AI. So both experiences are improving.” Improving documentation is the single highest-leverage action on this diagram.
The Neo4j question
Neo4j is confirmed for Kiki, the internal assistant. Only Sequoia's write-up extends it to the customer-facing path. Klarna has never said that. Drawn dashed.
Do not draw Salesforce
Many reconstructions put Salesforce Service Cloud on the handoff path. Klarna publicly shut down Salesforce CRM in 2025 and has never named a replacement. The destination system is genuinely unknown.
Component & agent view — the full request path
Everything from session open to escalation, with the intent-classification gap marked honestly. Click any box for the exact quote it rests on. Use hide inferred to see only what Klarna, OpenAI, LangChain or Neo4j have actually stated.
The intent-classifier gap
Klarna names no intent classifier. The only routing evidence is LangChain's “routed requests and handled different tasks using the LangGraph framework.” Routing may well be LLM-driven inside the graph rather than a separate classification hop. Both boxes are drawn; the classifier is dashed.
The gateway is real
Klarna's own job advert names it: “You will manage and operate Klarna's AI Gateway, ensuring reliable, low-latency access to AI models at scale across the organisation… onboard teams to the AI Gateway, configure model access.” Stack: Python, FastAPI, Kubernetes.
Guardrails: tested, not documented
Orosz could not induce hallucination in 15 minutes of probing. But Colin Fraser did succeed with prompt injection, getting it to generate code. The guardrail implementation is entirely unpublished.
Critical path — “I want a refund”
End to end, with every hop tagged by evidence strength. Klarna has published zero function, tool or API names — the one call name shown is explicitly marked illustrative rather than left out, so you can see where the boundary of knowledge actually falls.
Metrics timeline — every primary source
The automation rate went up through the 2025 “AI reversal.” That is the most counterintuitive fact in this whole case study.
| Date | Source | Agent-equiv | % of chats | Conversations | Savings |
|---|---|---|---|---|---|
| 2024-02-27 | Press release | 700 FTE | two-thirds | 2.3M (first month) | $40M est. |
| 2025-03-14 | Form F-1 | >800 FTE | 62% | 20.1M | ~$39M (2024) |
| 2025-05-21 | Form F-1/A | >700 revised down | 66% | 23.1M | $39M |
| 2025-09-10 | 424B4 prospectus | >700 | 69% | 25.3M | $39M |
| 2025-11-18 | Q3'25 deck, slide 33 | 853 FTE | 81% | 28M/yr | $58M |
| 2026-02-26 | 20-F FY2025 | >850 | 80% | 31M since launch | ~$59M |
| 2026-05 / 2026-08 | Q1'26, Q2'26 | absent — AI CS metrics vanished entirely from investor materials | |||
Two source conflicts worth flagging
(a) The F-1 said “>800” in Mar 2025, then the F-1/A said “>700” in May 2025 for a later period, unexplained. (b) LangChain's “2.5 million conversations to date” (Feb 2025) is irreconcilable with Klarna's own 20.1M one month later. Treat LangChain's count as unreliable; its architecture claims are sound.
“700 agents” was never a headcount
The 20-F is explicit: it is “estimated based on the average monthly reduction in chat and telephone conversations handled by full-time agents.” And those agents worked for outsourcing partners, not Klarna.
Confirmed vs inferred — the honest boundary
| Component | Status | Evidence |
|---|---|---|
| LangGraph orchestration/routing | confirmed | LangChain case study carrying a Siemiatkowski quote: “routed requests and handled different tasks using the LangGraph framework” |
| LangSmith tracing + LLM-as-judge | confirmed | Same source, plus independently corroborated by a Klarna GenAI job advert |
| RAG | confirmed | CEO on the Sequoia podcast, on the record |
| OpenAI as the model provider | confirmed | Press release, OpenAI case study, 20-F. Model version: never confirmed by anyone. |
| Klarna AI Gateway | confirmed | Klarna's own job posting names it and describes its function |
| Authenticated per-order context injection | confirmed | Observed empirically in a hands-on test |
| Human handoff on out-of-scope topics | confirmed | Observed empirically + 20-F “dual-track approach” |
| Refunds / returns / payments in scope | confirmed | OpenAI story, LangChain, 20-F all agree |
| Mandatory AI disclosure to customer | confirmed | CEO, Sequoia podcast |
| GenAI excluded from credit underwriting | confirmed | 20-F MD&A, verbatim |
| Discrete intent classifier | inferred | Never named. Routing may be LLM-internal to the graph. |
| Vector database product | inferred | No product ever named for the CS path |
| PII scrubbing before OpenAI calls | inferred | “We have — of course — solved for the data management aspects” is the entire record |
| Guardrail implementation | inferred | The teardown's author explicitly flags his own read as assumption |
| Neo4j on the customer-facing path | inferred | Confirmed for Kiki; CS link is Sequoia's characterisation only |
| Contact-centre platform for handoff | unknown | Salesforce was shut down in 2025; no replacement named |
| Any function / tool / API names | none exist | Zero published. Treat every named call in any Klarna diagram as invented. |
The “reversal” — four architectural deltas
Klarna did not de-automate. Deflection rose 62% → 80% across exactly the period the press described as a retreat from AI. What changed was the egress policy and the human supply model.
| # | Delta | What it means on the diagram |
|---|---|---|
| 1 | Unconditional human egress becomes an invariant | Pre-2025 handoff fired on guardrail trip or out-of-scope. The 20-F now codifies a “dual-track approach” offered to all customers — a permanent user-initiated escalation edge running parallel to the model-triggered one. |
| 2 | Human agent supply layer replaced | Out: ~3,000 BPO contractors. In: directly-recruited remote agents sourced from the customer base — the “Uber model.” Implies a new marketplace / scheduling / identity layer and access control for non-employees touching financial data. All undisclosed. |
| 3 | Segment-aware routing | AI positioned as commodity tier-1; human deliberately positioned as VIP. Implies value/segment routing layered on top of intent routing. inferred |
| 4 | Disclosure went dark | Zero standalone AI customer-service press releases across 135 investor news items in 2025–26. Metrics moved into filings only, then vanished from Q1'26 and Q2'26 entirely. |
Sources
- Klarna press release — AI assistant handles two-thirds of chats in its first month (27 Feb 2024)
- Klarna Group plc Form 20-F, FY2025 (SEC, 26 Feb 2026) — deflection rate, dual-track approach, underwriting exclusion, AI risk factors
- Q3'25 investor presentation, slide 33 — the 853 / 81% / $58M figures
- 424B4 IPO prospectus (10 Sep 2025) — chat categorisation tool, credit-denial explainer
- LangChain customer story: Klarna (12 Feb 2025) — LangGraph, LangSmith, dynamic prompting key technical source
- OpenAI customer story: Klarna
- Neo4j customer story: Klarna — Kiki and the knowledge graph
- Sequoia Training Data podcast with Sebastian Siemiatkowski (23 Jul 2024) — RAG, Neo4j, semantic search, the AI-disclosure rule, “secret sauce” key technical source
- Sequoia — Better agents need better documentation
- The Pragmatic Engineer — Klarna's AI chatbot: how good is it? — independent hands-on teardown key technical source
- CX Dive — Klarna reinvests in human talent (9 May 2025)
- CX Dive — Klarna pursues Uber-style customer service model (20 Feb 2026)
- CX Today — what actually replaced Salesforce at Klarna
