Loading repository data…
Loading repository data…
voidful / repository
Unlock the Power of LLM: Explore These Datasets to Train Your Own ChatGPT!
A transparent discovery signal based on current public GitHub metadata.
This score does not audit code, security, maintainers, documentation quality, or suitability. Verify the repository and its current documentation before adoption.

git clone https://github.com/voidful/awesome-chatgpt-dataset.git cd awesome-chatgpt-dataset/mixed/datasetpick whatever dataset you want to use, then merge and upload:
python preprocess.py your_dataset_name_to_HuggingFaceHub
Sorted by dataset size (small → large). Items with unknown size appear at the end.
| Dataset Name | Size | Languages | Source | License |
|---|---|---|---|---|
| TheoremQA | 1K | English | We annotated 800 QA pairs covering 350+ theorems spanning across Math, EE&CS, Physics and Finance. | mit |
| LIMA | 1K | English | LIMA: Less Is More for Alignment. | cc-by-nc-sa-4.0 |
| WildGuardMix | 1.7K | English | Safety training mixture with vanilla/adversarial prompts and multi-annotator labels. | odc-by |
| Berkeley Function Calling Leaderboard (BFCL) | 2K | English + code | Function/tool-calling eval covering parallel/multi-call scenarios across languages. | - |
| im-feeling-curious | 3K | English | Extract from Google’s “I’m Feeling Curious” facts. | - |
| Puffin | 3K | English | Exactly 3,000 multi-turn examples; each response via GPT‑4. | apache-2.0 |
| cc_sbu_align | 4K | English | MiniGPT‑4 alignment data (image–text). | bsd-3-clause |
| QA-Feedback | 4K | English | Re‑constructed ASQA with human feedback. | - |
| SLF5K | 5K | English | Summarization with Language Feedback (5K unique samples). | apache-2.0 |
| blended_skill_talk |
| 7K |
| English |
| 7k conversations blending personality, empathy, and knowledge. |
| - |
| GSM‑IC | 8K | English | Grade‑School Math with Irrelevant Context (distractor sentences). | - |
| ChatAlpaca‑10K | 10K | English | 10,000 multi‑turn conversations (Alpaca‑based). | apache-2.0 |
| PKU‑SafeRLHF‑10K | 10K | English | First‑round Safe‑RLHF data with safety preferences. | - |
| Dolly‑15K | 15K | English | 15k instruction records crowdsourced by Databricks. | cc-by-3.0 |
| WebGPT (comparisons) | 20K | English | Human preference comparisons for WebGPT reward modeling. | - |
| CodeAlpaca‑20K | 20K | English | 20,022 instruction–code pairs for code generation. | - |
| HelpSteer2 | 21K | English | Open-source helpfulness data for reward models and preference learning. | cc-by-4.0 |
| openapi-function-invocations‑25k | 25K | English | Synthetic + extracted OpenAPI function-call traces. | mit |
| LongForm | 28K | English | Reverse‑instruction long‑text generation dataset. | mit |
| Chatbot Arena Conversations | 33K | English | 33K cleaned Arena chats with pairwise preferences. | - |
| HC3 | 37K | English, Chinese | 37,175 instructions with human vs LLM answers. | - |
| Anthropic HH Golden | 45K | English | Helpful & Harmless preference data; golden subset. | - |
| Mol‑Instructions | 48K | English | Biomolecular instruction dataset for LLMs. | cc-by-4.0 |
| RefGPT | 50K | English, Chinese | Cost‑effective pipeline to generate multi‑turn Q&A with references. | - |
| arxiv‑math‑instruct‑50k | 50K | English | QA pairs derived from arXiv math abstracts. | - |
| arxiv‑math‑instruct‑50k (ArtifactAI) | 51K | English | T5‑generated questions; GPT‑3.5 answers. | - |
| Traditional Chinese Alpaca | 52K | Traditional Chinese | Alpaca translated by ChatGPT API. | apache-2.0 |
| Cabrita Dataset | 52K | Portuguese | Alpaca translated to Portuguese. | - |
| Japanese Alpaca | 52K | Japanese | Alpaca translated by ChatGPT API. | cc-by-nc-4.0; OpenAI terms |
| Alpaca Dataset | 52K | English | 175 seed instructions completed by OpenAI. | cc-by-nc-4.0; OpenAI terms |
| Alpaca Data Cleaned | 52K | English | Cleaned Alpaca 52K. | - |
| Alpaca GPT‑4 Data | 52K | English | Same prompts, GPT‑4 completions. | - |
| Alpaca GPT‑4 Chinese | 52K | Chinese | GPT‑4 completions for Chinese prompts. | - |
| xLAM Function Calling 60K | 60K | English | Structured tool-calling data for executable agents. | apache-2.0 |
| Dynosaur | 66K | English | Dynamic growth paradigm for instruction curation. | apache-2.0 |
| Finance | 69K | English | 68,912 finance‑related instructions. | - |
| WizardLM evol | 70K | English | Evolutionary instruction tuning data (WizardLM). | - |
| Vicuna Dataset | 75K | English | ~100k ShareGPT conversations (curated). | - |
| InstructionTranslation | 80K | Multi-lingual | M2M‑12B translated instructions (≤512 tokens). | mit |
| Self‑Instruct | 82K | English | 52K seed instructions; 82K I/O pairs. | - |
| OASST1 | 89K | Multi-lingual | Human‑generated assistant conversations (35 languages). | apache-2.0 |
| HH‑RLHF | 91K | English | Helpful/harmless RLHF pairs. | mit |
| Guanaco Dataset | 98K | En, Zh‑CN, Zh‑HK/TW, Ja | 175 Alpaca tasks across languages. | gpl-3.0 |
| InstructionWild | 104K | English, Chinese | Seeded 429 instructions; ~52K generated. | research-only; OpenAI terms |
| CAMEL Dataset | 107K | English | Multi‑role, topic‑diverse instruction dialogues. | - |
| TAPIR‑Cleaned | 117K | English | Cleaned IFTTT rule dataset for instruction tuning. | cc-by-nc-4.0 |
| OASST2 (final) | 135K | Multi-lingual | Open Assistant Conversations Release 2 (train+val). | apache-2.0 |
| WizardLM Evol‑Instruct V2 | 143K | English | 143K mixture‑evolved data. | - |
| LLaVA Visual Instruct 150K | 150K | English | GPT‑generated multimodal instruction pairs. | cc-by-nc-4.0 |
| ProsocialDialog | 166K | English | 165,681 prosocial instructions and feedback. | - |
| M2Lingual | 175K | Multi-lingual | Multilingual mixed‑modal (code+text) chat/instruct SFT. | - |
| COIG | 191K | Chinese | Chinese Open Instruction Generalist. | apache-2.0 |
| orca‑chat | 198K | English | Cleaned, pruned conversation‑style Orca subset. | - |
| OpenR1‑Math‑220k | 220K | English | DeepSeek‑R1 distilled math traces (verified). | apache-2.0 |
| Unnatural Instructions | 241K | English | Large creative/diverse instruction corpus. | mit |
| WildJailbreak | 262K | English | Synthetic jailbreak and benign contrastive prompts. | odc-by |
| SHP | 358K | English | 385K Reddit preference pairs across 18 topics. | reddit – revocable, non‑exclusive |
| Dromedary | 361K | English | Dromedary‑Verbose‑Clone synthetic instructions. | cc-by-nc-4.0 |
| UltraChat | 404K | English | Dual‑API generation (user vs assistant) for quality control. | cc-by-nc-4.0 |
| IGN Clean Instruct 500K | 509K | English | ~508k Ultrachat‑sourced, high‑quality instructions. | apache-2.0 |
| ELI5 | 559K | English | Long‑form community Q&A (“Explain Like I’m Five”). | - |
| GPT4All | 806K | Multi-lingual | LAION OIG + StackOverflow + P3 prompts; OpenAI outputs. | - |
| Instruct | 889K | English | 888,969 English instructions (augmented). | mit |
| MOSS | 1M | Chinese | GPT‑3.5‑turbo generated Chinese SFT data. | apache-2.0 + agpl-3.0 |
| WildChat | 1.0M | English | In‑the‑wild user–LLM chat dataset (license updated). | odc-by |
| smolTalk | 1.1M | English | Ultra‑compact multi‑turn chat for small‑scale SFT. | apache-2.0 |
| Open‑PerfectBlend | 1.42M | English | Diverse, deduped chat blend for general SFT. | apache-2.0 |
| The Tome | 1.75M | English | Large cleaned instruction dataset curated by Arcee. | mit |
| NaturalReasoning | 2.8M | English | 2.8M challenging reasoning questions (decontaminated). | cc-by-nc-4.0 |
| LaMini‑Instruction | 3.0M | English | ~2.58M–3M instruction–response pairs (GPT‑3.5). | cc-by-nc-4.0 |
| OpenOrca (full) | 3.0M | English | GPT‑4/3.5 augmented FLAN collection. | - |
| WildChat‑4.8M (nontoxic subset) | 3.20M | English | Nontoxic filtered split of WildChat 4.8M. | odc-by |
| Infinity‑Instruct | 8.9M | Multi-lingual | 7.4M base + ~1.5M chat instruction data. | cc-by-sa-4.0 |
| [BELLE‑10M](ht |