Athletes

SAP RealSpend: Anomaly Detection

UX Design, AI, Research

_

© 2017 Q4 – 2018 Q1

Nanda Dias | Staff Designer

© 2026

Nanda Dias | Staff Designer

SAP RealSpend: designing machine learning to drive adoption

SAP RealSpend: designing machine learning to drive adoption

Not every problem machine learning touches is one it should own. Screened from dozens of ideas down to the one problem only ML could genuinely solve, Anomaly Detection shipped and was featured live at SAP Sapphire, before the GenAI hype.

The Challenge

Managers approving cost center spending weren't trained in finance, so the real workflow was to find a controller and ask. That dependency was the actual problem: SAP wanted managers using RealSpend directly. We had three months to find out whether machine learning could simplify that workflow, starting from a genuinely open brief.


As Senior UX Designer, I worked alongside a Product Owner, a junior visual designer, and backend engineers evaluating ML feasibility in parallel.


Before any concept had a name, the team sketched broadly and separately. I facilitated a design thinking workshop with developers, our PO and PM, narrowing the field to three concepts: cost center anomaly detection, expense forecasting, and smart tagging on line items, each kept at paper-prototype fidelity so no one got attached before the data had a say. It's a technique I'd apply differently today, but it was Lean UX in practice: validate cheap, move fast. It was the first time our engineers had built with ML, and the first time our users had seen ML-generated content at all.

Decisions and Actions

  • Mapped the budget-approval ecosystem, managers, controllers, org chiefs, service providers, before writing a single problem statement, so the design targeted a real dependency.

  • Ran the exploration as a structured ideation workshop, scoring every idea against four criteria to do the narrowing.

  • Built paper prototypes with the engineers directly, so feasibility debates and design decisions happened in the same room, in real time.

  • Designed and moderated three rounds of validation, 28 sessions in total: two internal rounds with the same 10 experts, then external testing at SAP Sapphire / ASUG 2018, using a written moderator/note-taker script so data stayed comparable round to round.

  • Proposed running SUS scoring quarterly, giving the team a comparable usability baseline going forward.

  • Presented synthesised findings back to the team, including direct engineer feedback on whether the sessions changed how they worked.

The Challenge

Managers approving cost center spending weren't trained in finance, so the real workflow was to find a controller and ask. That dependency was the actual problem: SAP wanted managers using RealSpend directly. We had three months to find out whether machine learning could simplify that workflow, starting from a genuinely open brief.


As Senior UX Designer, I worked alongside a Product Owner, a junior visual designer, and backend engineers evaluating ML feasibility in parallel.


Before any concept had a name, the team sketched broadly and separately. I facilitated a design thinking workshop with developers, our PO and PM, narrowing the field to three concepts: cost center anomaly detection, expense forecasting, and smart tagging on line items, each kept at paper-prototype fidelity so no one got attached before the data had a say. It's a technique I'd apply differently today, but it was Lean UX in practice: validate cheap, move fast. It was the first time our engineers had built with ML, and the first time our users had seen ML-generated content at all.

Men Red BG
Men Red BG

Decisions and Actions

  • Mapped the budget-approval ecosystem, managers, controllers, org chiefs, service providers, before writing a single problem statement, so the design targeted a real dependency.

  • Ran the exploration as a structured ideation workshop, scoring every idea against four criteria to do the narrowing.

  • Built paper prototypes with the engineers directly, so feasibility debates and design decisions happened in the same room, in real time.

  • Designed and moderated three rounds of validation, 28 sessions in total: two internal rounds with the same 10 experts, then external testing at SAP Sapphire / ASUG 2018, using a written moderator/note-taker script so data stayed comparable round to round.

  • Proposed running SUS scoring quarterly, giving the team a comparable usability baseline going forward.

  • Presented synthesised findings back to the team, including direct engineer feedback on whether the sessions changed how they worked.

Men Red BG
Woman Greyscale
Woman Greyscale
Woman Greyscale

Design Iterations & Learnings

Certainty, redrawn

We designed the chart assuming flags were simply right or wrong: solid mustard fill on anything flagged, reading as certain. Validation proved otherwise, machine learning anomaly detection doesn't always return true positives, and the chart was overstating a certainty nobody had. We redrew it as a thin line: a hint to investigate, not a verdict to trust. We also tested surfacing the algorithm's calculation logic to build trust further, but it overcomplicated the UI. Instead, clicking a flag opens the underlying line item itself, so managers can judge for themselves without needing to understand the model.

Speaking the manager's language

Early feedback cut deeper than the chart itself: "This app wasn't thought out to meet managers' vocabulary or way of thinking. It feels focused on controllers instead." We rebuilt the chart's language around the manager. "Actual" became "Spent," a word managers already used. We dropped "Committed" entirely, redundant next to "Approved" for this audience. And we moved the anomaly indicator off mustard, a colour that reads as danger, onto a neutral tone in the product's own palette, so a flagged expense reads as worth a look.

Making the toggle visible

The detection toggle had the same visibility problem from a different angle: it lived below the chart, an area validation showed people barely noticed. It was important enough that people wanted to see it at a glance, so we moved it into the header, top right, next to the other primary controls.

Reporting, moved inline

The first version let managers flag a false anomaly back to the system over conventional email. It worked, technically, but validation showed people didn't want to leave the tool to do it. Three rounds of iteration moved reporting into the in-app digital assistant, with a confirmation that the report had been read. External experts at SAP Sapphire confirmed people wanted to investigate inline.

Cutting scope on the data

Smart Tagging tested weaker than the other two concepts on both desirability and feasibility, so we dropped it. Forecasting ranked higher on perceived value; Anomaly Detection came second. But the feasibility debate surfaced the deciding fact: Forecasting could plausibly be solved with non-ML statistical methods, weakening the case for spending ML effort there. Anomaly Detection was the harder problem, and the one only machine learning could genuinely address. The call came down to technical fit.

Reflections

Not everyone believed this would work. The loudest concern going in: over-alert, and managers would tune the feature out entirely, one more distraction from the job they were actually there to do.

A deeper risk sat underneath that one. The model worked probabilistically, comparing current spend against the same period in previous years, a baseline that broke down whenever a team had gone through a reorg. Getting it right depended on user feedback, and five requests in validation asked for exactly the failure mode I was worried about: let me dismiss an anomaly without training the algorithm on it.

A flat dismiss isn't enough. "Not an anomaly" can mean a one-timer, or it can mean stop flagging this pattern altogether, and collapsing both into a single button risks teaching the model to stop looking, even in the year it shouldn't.

Either way, the system still needed a signal. I pushed the team to design feedback options that captured which kind of "no" it was actually getting.

Good team collaboration builds momentum and saves rework: a shared understanding of the problem, and of where a feature sits in the bigger picture, before jumping to solutions. That pays off even faster now, when misalignment compounds quickly once building starts. Paper prototyping specifically feels dated today, but the discipline behind it doesn't. Team collaboration and DesignOps have always mattered to me, and this project is a strength I want to keep carrying forward as roles keep blurring.

design iterated
Man Dancing

Evolving as a team

testimonials

workshop

Raja A.

Director Product Development, Wikimedia Deutschland

Raja A.

Director Product Development, Wikimedia Deutschland

Raja A.

Director Product Development, Wikimedia Deutschland

Questions people usually ask

Expand all

(01)

Vellfire Calibration

Art Direction

© 2025

What does complex, large-enterprise B2B SaaS experience actually give me, and would I still move fast enough for a smaller team?

How do I work when the brief isn't clear yet, and what part of the job do I enjoy most?

How do I actually strengthen a team, and will I get hands-on or mostly direct from a distance?

What did I do with the time between roles, where do I stand on AI, and what am I looking for next?

How has living abroad shaped me, both as a person and as a principal design lead?

Nanda Dias | Staff Designer

© 2026

Nanda Dias | Staff Designer

more works