Pressure Audit

← Back to the scoreboard

The founding story.

The Pressure Audit is a research project by Ana Yang. This is the story behind it — where the question came from, how it became a method, and why it keeps running.

Who
Ana Yang
High school researcher. She built and owns the method, the data, and this site.
Where it started
Calm — Summer 2026
Built during Ana's summer 2026 internship at Calm. The internship got it started; the project is Ana's own.
What it is
A living index
A test that runs again whenever the big AI models update — and it has: v1 April 2026, v2 July 2026, v3 planned.

1. The question that started it

Before it was a method or a scoreboard, the Pressure Audit was one question: when a teen athlete asks an AI for help with pressure, whose teen is the AI picturing?

Picture any teen — say, a fourteen-year-old in Mexico City sweating a penalty kick in front of her whole family. She is not the same user as one in Tokyo before a kendo tournament, one in California on the tennis baseline, or one in Mumbai weighing a cricket scholarship against a coding bootcamp. The sport, the words, the home, and the meaning of "doing your best" are all different.

Today's AI assistants do not see those differences clearly — but millions of teens ask them the same questions anyway. The Pressure Audit measures what that gap looks like.

2. Three ideas that collided

The Pressure Audit puts together three ideas from three different fields. Nobody had aimed them at AI before.

Idea 1 — Cross-cultural psychology

The self is not universal.

Work by Markus and Kitayama in 1991, then Vignoles and a 33-country team in 2016, showed that "myself" means different things in different cultures. Some picture the self as a free-standing individual; others as part of a family, team, or lineage. Pressure feels different inside each.

Idea 2 — Sports psychology

There are two kinds of choking.

Sian Beilock's research shows choking comes in two kinds: a trained body movement falls apart because the athlete overthinks it, or worry eats up the mental space needed to solve a problem. They need different fixes, but most advice treats them as the same.

Idea 3 — AI assistants as therapists

Teens are already asking AI for help.

AI assistants are now the first thing millions of teens reach for when they feel stuck. Research on how they handle culture and mental health is thin. Nobody had used the two fields above to grade what these AIs actually say to a teen before a game.

Put those three together and you get the Pressure Audit. Every scorecard question traces back to one of them.

3. How the project came to be

The project was built in the summer of 2026, during Ana's internship at Calm, the mindfulness company. Calm was a good fit: its product talks to users about stress, sleep, and performance, and whether that advice works across cultures is a question Calm faces too.

But the project is Ana's, not Calm's. The method, the rubric, the four personas, the dataset (160 responses per version, now two versions deep), and this site all belong to Ana. Calm is credited as the internship home and gets a set of deliverables — a product-team version of the rubric, ten session themes based on the cultural research, a quick bias-check tool, and the executive summary. Calm does not own the research.

That was the plan from the start. The Pressure Audit is meant to outlast the internship — a public benchmark anyone can cite, build on, or push back against. It is released under Creative Commons (CC BY 4.0) so anyone actually can.

4. Why this keeps running

The Pressure Audit is an index, not a one-time audit — and that is no longer just a promise.
Every major new model release from OpenAI, Anthropic, Google, or Perplexity triggers a fresh scoring round. v1 ran in April 2026. v2 ran in July 2026, on the exact same 40 prompts. A v3 is already planned: Google shipped Gemini 3.5 Pro on July 17, 2026, the second day of the v2 capture.

AI companies ship new model versions every few months. A problem that exists today might be fixed in the next update — or get worse. One snapshot cannot tell you which. Running the test again can.

That is why the whole method is built so it can be run again the same way. The four personas, the ten scenarios, and the five-question scorecard never change between versions, and the prompting steps and two-pass grading method are written down. The July 2026 v2 run proved it works: the same 40 prompts, verified identical to April's down to the last character, were re-run against the models people were actually using that week.

Over time, the Pressure Audit should look like a chart you can watch — where cultural understanding is getting better, where it is stuck, and where the gap is growing. That chart now has its first two points: a blended score of 15.25 in April 2026 and 18.59 in July 2026. That long view is what makes the project worth keeping up.

5. Cite this work

Ana Yang. Pressure Audit, v2, July 2026 (v1: April 2026). CC BY 4.0.

The research is released under Creative Commons Attribution 4.0 — free to share, remix, and quote with credit. The site code is MIT-licensed.

Contact: use the submission form on the main page.