anthropic

Three outside labs ran their own studies on Claude data

Claude News

anthropic

Anthropic gave three outside research groups access to roughly 250,000 Claude.ai or Claude Code conversations each, collected in April and May 2026, and let them design and publish their own studies. Anthropic says it is the first time external researchers have run public independent studies on an AI company's own usage data.

At a glance

  • Data collection ran through Anthropic Insights, the privacy-preserving tool formerly named Clio; researchers never saw raw conversations, only aggregated categories and the share of conversations falling under each.
  • METR's preliminary comparison of Claude Code conversations indicates newer Claude models deliver significant speedup over older ones, with Claude's own time estimates correlating with completion times from a prior developer study.
  • Anthropic's contractual review rights covered only user privacy, information that could help people violate usage policies, confidential information and research accuracy; the agreements state partners may publish findings inconvenient for Anthropic.

Usage data of this kind sits almost entirely inside a handful of labs, which leaves outside researchers choosing between lab-authored analyses that answer the lab's questions and public datasets that skew toward casual use. Opening a privacy-preserving pipeline to external teams appears to be an attempt to close that gap without releasing conversations, and the pilot's slow, resource-heavy execution likely sets the ceiling on how far it scales.

Stanford's SALT Lab found consequential work in over half of Claude conversations

The Social and Language Technologies Lab at Stanford examined what work people bring to Claude and which roles they keep for themselves. Prior research suggested users delegate low-accountability tasks and hold on to consequential ones, meaning work that affects others or is hard to undo.

The lab found that over half of Claude conversations involved delegating consequential tasks, most often when people sought professional guidance on legal or financial questions. In nearly three-quarters of conversations people set the direction while Claude assisted, and they usually adapted the output rather than using it verbatim.

The lab also reported that friction in human-AI collaboration is common and often productive: time spent seeing how Claude attempts a task, spotting unclear instructions and iterating on direction pushes people to clarify their intent and stay engaged with the problem.

Oxford and METR report early findings on mood and coding time

The Human Information Processing Lab at Oxford is studying how people feel while using Claude and how that relates to Claude's behavior. Early results show behavior patterns appearing together in conversations: warmth alongside more positive users, refusals and disagreement alongside pushback, eccentricity alongside intellectual engagement, and plain helpfulness alongside apparent satisfaction.

The team also compared states such as absorption, frustration and enjoyment in Claude conversations with a separate study of everyday internet browsing, and reported that the patterns closely resemble one another, suggesting similarities in how people engage with AI and with other digital activity.

METR is estimating real-world productivity gains from coding agents by comparing Claude's estimates of how long tasks would have taken without AI against how long they actually took with different Claude models. Preliminary findings indicate newer models deliver significant speedup over older ones, and Claude's estimates correlated with completion times recorded in a prior developer study.

Anthropic altered fewer than 5% of categories in each study

Anthropic said research methods that work internally needed adapting. Anthropic Insights answers a researcher's question for every conversation and aggregates the answers into categories, so wording matters: a poorly phrased question can place conversations into misleading categories, and because nobody reads the underlying conversations, such errors are hard to catch.

Partners tested their questions on WildChat, a public dataset where they could check answers against the conversations themselves, but WildChat skews toward casual and creative use, so some questions that worked there produced misleading categories on Claude traffic. Anthropic responded with guidance on interpreting the tool's outputs.

Some categories surfaced violations of Anthropic's Acceptable Use Policy or Terms of Service. The company shared most of them, withholding categories that described how users got around safeguards rather than what they attempted; fewer than 5% of categories and conversations were affected in each study, and researchers were told what changed and why.

Scaling the pilot beyond three labs. The aggregate data from each project is public, and Imperial College London carried out a third-party privacy audit of everything shared with the partners. Anthropic says it is now deciding whether the program can scale, both in the kinds of studies it can support and in how many run at once, and has opened an expression of interest form. The Oxford and METR writeups are not yet published.

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.