the catalog's drawers
Sources & licenses
ChatCorpus holds 5M+ conversations from ten research dataset releases. Source datasets applied PII redaction before publication; content is shown unmodified apart from length truncation. Conversations flagged by the source datasets' moderation tooling are labeled and can be excluded with the content filter. To request removal of a conversation, contact us with its record number.
3.2M real user conversations with ChatGPT (GPT-3.5 and GPT-4), collected by Ai2 with opt-in consent from April 2023 to May 2024, 74 languages, deduplicated across releases. Attribution: Zhao et al., "WildChat: 1M ChatGPT Interaction Logs in the Wild" (2024).
1M real conversations with 25 models across 154 languages from the LMSYS Chatbot Arena platform (2023). Included under license. Attribution: Zheng et al., "LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset" (2023).
~660k conversations with human preference votes across four releases: the original 33k pairs (2023), 55k (Apache-2.0, 2024), 100k (2024) and 140k (2025) — spanning 70+ models from GPT-4 and Claude to Gemini and Llama, with the human vote on each pair. Attribution: LMSYS / LMArena.
24k battles between search-augmented models (Perplexity-style assistants), March–May 2025, ~90 languages. Attribution: LMArena (2025).
111k crowdsourced assistant conversations in 28+ languages, linearized from message trees; humans wrote both sides. Attribution: Köpf et al., "OpenAssistant Conversations" (2023).
161k human preference pairs (chosen vs. rejected assistant responses) from Anthropic's helpfulness and harmlessness research. Attribution: Bai et al., "Training a Helpful and Harmless Assistant" (2022).
Alongside these open sources, we're also building our own conversation datasets — collected and curated by us, and rolling out to ChatCorpus over time.