Loupecase studyDashboard →

Opportunity brief

What this is for: Make the case for building Loupe before any design or code work starts. Date: 2026-09-20 Status: Final

Problem

A person who owns an AI assistant feature cannot answer who uses it, how intensely, for what, or where it fails, without an engineer writing a one-off script. Priya is the product manager for the AI assistant inside a customer-support tool, and about 40,000 conversations a week flow through it. Her VP asks three questions on a Monday: is usage growing or is it the same power users every week, what are people using it for, and support tickets say the assistant "doesn't get it" so where is it failing. Priya asks a data engineer for a SQL pull, waits three days, gets a spreadsheet of 2,000 conversations, and reads a few hundred by hand over a weekend to form a gut impression that nobody else can check. A month later the same questions come back and she repeats the whole cycle.

Who has it

Evidence

Incumbent tools (DeepEval, Braintrust, LangSmith, promptfoo, Confident AI) are developer tools that assume the user already knows their metrics and writes code. WildVis, the dataset authors' own tool, is a search-and-browse visualizer for researchers, not a metrics product for product owners. Loupe is demonstrated on WildChat-4.8M, 3.2 million real human-ChatGPT conversations, standing in for a team's own production logs. Its findings describe this dataset's population; they are suggestive, not representative, of AI assistant users in general.

Why now

35% of teams shipping AI still run no evaluation at all, and choosing what to measure, not running the measurement, is the hard part. Evaluation tooling matured fast over the last two years, and it matured for the people who write the prompts and the code, not for the people who own the product decision. A PM who is accountable for an assistant's quality and adoption still has no product built for her.

The wedge

Loupe is opinionated, PM-facing, metrics-first analytics over conversation logs, not a general eval framework. It covers five families a product owner asks about, in this order: volume (who and how many), intensity (how often, growing or decaying), intent (what people are trying to do), friction (where the assistant fails), and data quality (how much to trust the rest). Friction, not a full eval product, is the wedge. A full eval product asks an engineer to define pass/fail criteria for every response, which recreates the same bottleneck this brief opens with. Friction proxies (a repeated request, a one-turn abandonment, a correction, a refusal pattern) are computed from conversation structure alone, need no code from the team, and point a PM straight at which failure mode to escalate. That claim is narrower than "we score every response," and it is one a PM without an eval engineer can use.

What we will not do

Decision requested

Build v1 as specified in the PRD, demonstrated on WildChat, with all findings labeled as describing that dataset's population, time-boxed to eight working days.