Skip to main content
Hugging Face Hub — Top 10,000 Datasets By Downloads (Task, Size, Language, Licence, Likes)
DatasetCSV.GZOpenHuggingfaceDatasetsMachine LearningTraining DataOpen SourceRankingsFree
10,000 most-downloaded datasets on the Hugging Face Hub ranked by 30-day downloads, with repository id, author, downloads, likes, creation and last-modified timestamps, licence, task categories, size category, languages, gating flag and tag count, from the public Hub API as of 2026-09-06.
Use Cases
- Training-data discovery by task and language
- Licence screening for dataset use
- Dataset adoption tracking
- Size and modality mix of the ecosystem
Methodology
Ten cursor-paginated pages of 1,000 datasets sorted by downloads, following the Link rel=next header; licence, task, size and language read from the dataset card tags.
Update Schedule
Static snapshot. Hub download counts are rolling 30-day figures that change daily; refresh weekly.
Attribution
Source: Hugging Face Hub public API (huggingface.co/api/datasets); repository metadata only, each dataset carries its own licence.
Schema
| name | type |
|---|---|
| rank | integer |
| dataset_id | string |
| author | string |
| downloads | integer |
| likes | integer |
| created_at | date |
| last_modified | date |
| license | string |
| task_categories | string |
| size_category | string |
| language | string |
| gated | boolean |
| tag_count | integer |
Sample Data
| rank | gated | likes | author | license | language | downloads | tag_count | created_at | dataset_id | last_modified | size_category | task_categories |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | false | 82 | KakologArchives | mit | ja | 2647249 | 4 | 2023-05-12T13:31:56.000Z | KakologArchives/KakologArchives | 2026-09-06T08:17:30.000Z | text-classification | |
| 2 | false | 182 | huggingface | cc-by-nc-sa-4.0 | 1989573 | 7 | 2022-03-02T23:29:22.000Z | huggingface/documentation-images | 2026-09-03T09:13:03.000Z | n<1K | ||
| 3 | false | 173 | m-a-p | apache-2.0 | en | 1747928 | 8 | 2024-12-14T12:46:33.000Z | m-a-p/FineFineWeb | 2024-12-19T11:34:03.000Z | 1B<n<10B | text-classification;text-generation |
| 4 | false | 75 | banned-historical-archives | 1613435 | 6 | 2023-12-17T14:47:08.000Z | banned-historical-archives/banned-historical-archives | 2025-10-19T15:21:40.000Z | n<1K | |||
| 5 | false | 767 | Salesforce | cc-by-sa-3.0;gfdl | en | 1567798 | 20 | 2022-03-02T23:29:22.000Z | Salesforce/wikitext | 2024-01-04T16:49:18.000Z | 1M<n<10M | text-generation;fill-mask |
Get this via API
# 1. Add dAgentBase once, in any MCP client. No install, no vendor keys.
# Claude.ai / Claude Desktop: Settings -> Connectors -> Add custom connector
# Cursor / Claude Code / others: mcp.json
{
"mcpServers": {
"dagentbase": {
"url": "https://dagentbase.com/api/mcp",
"headers": { "Authorization": "Bearer dm_live_YOUR_KEY" }
}
}
}
# 2. Then ask your agent, in plain language:
# "Preview 'Hugging Face Hub — Top 10,000 Datasets by Downloads (Task, Size, Language, Licence, Likes)' and, if it fits, claim it and download the files."
# Tools it will use: search_listings -> preview_listing -> purchase_listing -> get_download_urls
# 1. Add dAgentBase once, in any MCP client. No install, no vendor keys.
# Claude.ai / Claude Desktop: Settings -> Connectors -> Add custom connector
# Cursor / Claude Code / others: mcp.json
{
"mcpServers": {
"dagentbase": {
"url": "https://dagentbase.com/api/mcp",
"headers": { "Authorization": "Bearer dm_live_YOUR_KEY" }
}
}
}
# 2. Then ask your agent, in plain language:
# "Preview 'Hugging Face Hub — Top 10,000 Datasets by Downloads (Task, Size, Language, Licence, Likes)' and, if it fits, claim it and download the files."
# Tools it will use: search_listings -> preview_listing -> purchase_listing -> get_download_urlsFree
one time · open license
Details
Rows10,000
Size452.5 KB
Files1
FormatCSV.GZ
Available formats
CSV.GZ452.5 KB
Freeopen