AI Observatory:独立追踪真实AI使用数据
We still don’t know how people are really using AI
AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but they only release the data they want us to see, AI researchers say.
像Anthropic和OpenAI这样的人工智能公司定期发布关于人们如何使用Claude和ChatGPT等产品的报告,但人工智能研究人员表示,他们只发布他们希望我们看到的数据。
“There is no independent source to corroborate it,” says Anka Reuel, a Computer Science PhD candidate at the Stanford Trustworthy AI Research (STAIR) Lab.
“没有独立的来源可以证实这一点,”斯坦福可信人工智能研究(STAIR)实验室的计算机科学博士生Anka Reuel说。
Reuel is co-lead of a new research project, called the AI Observatory, that aims to fill in the gap. It’s a public platform that aggregated and analyzed real AI conversations with popular models like Claude and Gemini that were collected with users’ consent through seven existing datasets. The Observatory’s intent is to provide independent sources of information for researchers and policymakers to assess how people are using generative AI. Stakeholders are currently making highly consequential decisions about AI’s benefits and risks based on very limited data, says Reuel.
Reuel是名为AI Observatory的新研究项目的联合负责人,该项目旨在填补这一空白。这是一个公共平台,通过七个现有数据集,在用户同意的情况下,汇总并分析了与Claude和Gemini等流行模型的真实AI对话。该观察站的目的是为研究人员和政策制定者提供独立的信息来源,以评估人们如何使用生成式AI。Reuel表示,利益相关者目前正基于非常有限的数据,对AI的益处和风险做出影响深远的决策。
The AI Observatory found that AI use differs significantly across models, and has changed over time. Its research shows many more sensitive behaviors than are captured in reports from major AI companies, which they say focus more on work than on personal use.
AI Observatory发现,AI的使用在不同模型之间存在显著差异,并随时间发生了变化。其研究显示,许多敏感行为比主要AI公司报告中所捕捉到的更多,这些报告更侧重于工作而非个人使用。
Anthropic Economic Index is one of the best known and most widely cited sources of AI usage data but it has blind spots. As its name suggests, it focuses on work- and productivity-related uses of Claude AI—filtering out conversations that are unrelated to these uses.
Anthropic经济指数是最知名且被广泛引用的AI使用数据来源之一,但它存在盲点。正如其名称所示,它专注于Claude AI在工作与生产力相关的用途,过滤掉了与这些用途无关的对话。
When the AI Observatory team applied Anthropic’s methods to their dataset, they found that nearly half of the conversations—or 48%— would have been filtered out. Those non-work-related conversations that were filtered out were more likely to include health and relationships (44.2% versus 31.2% in Anthropic’s analysis), adult or illicit topics (7.9% versus 2.1%), harassment and hate (27.5% versus 5.66%), and sexual content (16.7% versus 2.4%). (OpenAI’s 2025 report on ChatGPT use, similarly, found that only 30% of consumer use was related to work.)
当AI Observatory团队将Anthropic的方法应用于他们的数据集时,他们发现近一半的对话(即48%)会被过滤掉。这些被过滤掉的与工作无关的对话更可能包含健康和人际关系(44.2%对比Anthropic分析中的31.2%)、成人或非法话题(7.9%对比2.1%)、骚扰和仇恨(27.5%对比5.66%)以及性内容(16.7%对比2.4%)。(OpenAI 2025年关于ChatGPT使用的报告同样发现,消费者使用中只有30%与工作相关。)
Anthropic has released separate blog posts on how people use Claude for support or companionship, and even to generate CSAM, “but having [the AI Observatory’s] bird eye view analysis rather than sectioned off into a separate report helps” researchers understand the different uses more consistently, says David Widder, an assistant professor at UT-Austin’s School of Information, who researches how people interact with AI systems and is not involved with the AI Observatory.
Anthropic已发布单独的博客文章,介绍人们如何使用Claude进行支持或陪伴,甚至生成CSAM,“但拥有[AI Observatory]的鸟瞰式分析,而不是将其分割成单独的报告,有助于”研究人员更一致地理解不同的用途,德克萨斯大学奥斯汀分校信息学院助理教授David Widder说,他研究人们如何与AI系统互动,并未参与AI Observatory项目。
The datasets the AI Observatory looked at include conversations that took place between 2023 and 2025, and it found differences both in how people were using AI and how various AI platforms responded.
AI观察站所审视的数据集涵盖了2023年至2025年间发生的对话,并发现了人们在如何使用AI以及不同AI平台如何回应方面的差异。
Conversations within WildChat, one of the largest and most detailed datasets included in the AI Observatory’s study, got longer and more elaborate over time, indicated by increases in prompt tokens, response tokens, and conversation turns.
在AI观察站研究包含的最大且最详细的数据集之一WildChat中,对话随时间变得更长、更详尽,表现为提示令牌、响应令牌和对话轮次的增加。
There was also significantly more small talk over time. That suggests that AI companionship was increasing; meanwhile the AI assistants’ self-disclosure (i.e. that it was a chatbot) decreased.
随着时间的推移,闲聊也显著增多。这表明AI陪伴感在增强;与此同时,AI助手的自我披露(即表明自己是聊天机器人)减少了。
Additionally, exchanges that the researchers labeled as sensitive use—meaning ones with potentially harmful or restricted content, including sexual harassment and hate speech—dropped. That might suggest that platforms were generally deploying more effective safeguards.
此外,研究人员标记为敏感使用(即可能含有有害或受限内容,包括性骚扰和仇恨言论)的交流减少了。这可能表明平台普遍部署了更有效的防护措施。
The AI Observatory also found that AI use looked significantly different depending on the model. Depending on the tool, users ranged in topics, interaction styles, conversation structures, as well as both the likelihood and type of sensitive use cases.
AI观察站还发现,AI的使用因模型不同而有显著差异。根据工具的不同,用户在话题、互动风格、对话结构以及敏感用例的可能性和类型上都有所不同。
For example, the researchers found that people used Grok and Gemini more frequently for information retrieval. Grok, in particular, was especially popular for information on news and politics, but it was also where misinformation tended to concentrate. (This is consistent with other research that has also shown how readily misinformation proliferates on Grok. xAI did not respond to a request for comment.)
例如,研究人员发现人们更频繁地使用Grok和Gemini进行信息检索。特别是Grok,在新闻和政治信息方面尤其受欢迎,但也是错误信息集中的地方。(这与其他研究一致,这些研究也显示了错误信息在Grok上易于扩散。xAI未回应置评请求。)
Meanwhile, people were more likely to turn to Anthropic for coding, Gemini for social and roleplay uses, and ChatGPT for homework assistance.
与此同时,人们更倾向于使用Anthropic进行编码,使用Gemini进行社交和角色扮演,以及使用ChatGPT进行作业帮助。
There were even differences among different versions of the same model. Researchers found that people had shorter conversations with ChatGPT when it was powered by GPT-3.5, and longer and more iterative ones with GPT-4o—which makes sense given that that version became known for leading to emotional addiction.
即使是同一模型的不同版本也存在差异。研究人员发现,当ChatGPT由GPT-3.5驱动时,人们的对话较短,而与GPT-4o的对话则更长且更具迭代性——考虑到该版本因导致情感上瘾而闻名,这很合理。
Company reports, however, didn’t tend to capture these nuances across or even within their own models. “No single company report tells the whole story,” says Shayne Longpre, a recent PhD graduate from the MIT Media Lab who co-led the research with Reuel.
然而,公司报告往往未能捕捉到这些跨模型甚至模型内部的细微差别。与Reuel共同领导这项研究的麻省理工学院媒体实验室近期博士毕业生Shayne Longpre表示:“没有哪一份公司报告能讲述完整的故事。”
To create the AI Observatory, Reuel and researchers from MIT, Stanford, the Data Provenance Initiative, and other institutions, aggregated 24,521 conservations across 85,633 conversational turns (that is, the user prompt and corresponding AI response) from seven real-world datasets collected by previous research. These conversations came from 5,000 users interacting with 52 different models, including ChatGPT, Gemini, Claude, and Grok, between 2023 and 2025.
为了创建AI观察站,Reuel与来自MIT、斯坦福、数据溯源倡议及其他机构的研究人员,汇总了先前研究收集的七个真实世界数据集中的24,521段对话,涵盖85,633个对话轮次(即用户提示及相应的AI回复)。这些对话来自2023年至2025年间,5000名用户与52种不同模型的互动,包括ChatGPT、Gemini、Claude和Grok。
But these conversations are a drop in the proverbial bucket compared to the data that the big labs themselves have access to. The latest Anthropic Economic AI Index, for example, is based on analysis of 1 million Claude conversations; OpenAI’s report on how people are using ChatGPT analyzed 1.5 million conversations.
但与大型实验室自身掌握的数据相比,这些对话不过是沧海一粟。例如,最新的Anthropic经济AI指数基于对100万段Claude对话的分析;OpenAI关于人们如何使用ChatGPT的报告则分析了150万段对话。
An Anthropic representative said that their published research reflects their research teams’ specific questions and interests, and the importance of supporting external independent research. OpenAI did not respond to requests for comment.
Anthropic的一位代表表示,他们发布的研究反映了其研究团队的具体问题和兴趣,以及支持外部独立研究的重要性。OpenAI未回应置评请求。
Additionally, the fact that the AI Observatory’s dataset draws from voluntarily-provided sources means that it’s likely underrepresenting sensitive uses, which people may be less likely to share. Thus, the researchers caution that its findings are not indicative of all AI use.
此外,AI观察站的数据集来源于自愿提供的资料,这意味着它可能低估了敏感用途,因为人们不太可能分享这类内容。因此,研究人员提醒,其发现并不能代表所有AI使用情况。
The Observatory’s work, though, broadens access for the research community. AI companies don’t typically share their chat data for analysis, which means their reports tend to focus on the findings that paint their companies in the best light, independent researchers, like Reuel and Widder, say.
尽管如此,观察站的工作拓宽了研究社区的访问渠道。独立研究人员如Reuel和Widder表示,AI公司通常不会分享其聊天数据供分析,这意味着他们的报告往往侧重于能展现公司最佳形象的发现。
“When we want to ask, for example: is Anthropic’s general-purpose AI system…used mostly for good, or mostly for bad…we don’t have a way of answering that question because that information is proprietary,” explains Widder, the assistant professor at UT-Austin’s School of Information.
“例如,当我们想问:Anthropic的通用AI系统……主要用于善还是恶……我们无法回答这个问题,因为该信息是专有的,”UT-Austin信息学院的助理教授Widder解释道。
The AI Observatory’s data will be available to researchers for analysis, and the team hopes to expand its datasets over time. Ideally, Reuel says, the AI companies would share their data with independent researchers—in ways that protect user privacy, of course. But as it currently stands anyone making decisions based on AI usage data risks “completely operating in the wild and making these really consequential decisions without knowing what’s actually happening beyond those company narratives,” says Reuel.
AI观察站的数据将提供给研究人员进行分析,团队希望随着时间推移扩大其数据集。Reuel表示,理想情况下,AI公司会以保护用户隐私的方式与独立研究人员分享数据。但就目前而言,任何基于AI使用数据做决策的人,都冒着“完全在未知领域操作,并在不了解公司叙事之外实际发生情况的情况下做出这些重大决策”的风险,Reuel说。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力