JavaScript is disabled. Some features may not work.
huggingface-datasets — ★ 10.7K GitHub Stars — Install Guide | SkillsNav
🇺🇸 English🇨🇳 中文
SkillsNav
Home

huggingface-datasets

★ 10K repogenerationN/AIntermediateClaude
🤖 AI Summary

Executes read-only Hugging Face Dataset Viewer API calls to explore, preview, paginate, search, filter, and retrieve metadata/parquet links from datasets, with support for gated/private datasets via token auth.

How to Install

Claude Code:
git clone --depth 1 https://github.com/huggingface/skills.git && cp skills/skills/huggingface-datasets ~/.claude/skills/huggingface-datasets -r
# Hugging Face Dataset Viewer Use this skill to execute read-only Dataset Viewer API calls for dataset exploration and extraction. ## Core workflow 1. Optionally validate dataset availability with `/is-valid`. 2. Resolve `config` + `split` with `/splits`. 3. Preview with `/first-rows`. 4. Paginate content with `/rows` using `offset` and `length` (max 100). 5. Use `/search` for text matching and `/filter` for row predicates. 6. Retrieve parquet links via `/parquet` and totals/metadata via `/size` and `/statistics`. ## Defaults - Base URL: `https://datasets-server.huggingface.co` - Default API method: `GET` - Query params should be URL-encoded. - `offset` is 0-based. - `length` max is usually `100` for row-like endpoints. - Gated/private datasets require `Authorization: Bearer `. ## Dataset Viewer - `Validate dataset`: `/is-valid?dataset=` - `List subsets and splits`: `/splits?dataset=` - `Preview first rows`: `/first-rows?dataset=&config=&split=` - `Paginate rows`: `/rows?dataset=&config=&split=&offset=&length=` - `Search text`: `/search?dataset=&config=&split=&query=&offset=&length=` - `Filter with predicates`: `/filter?dataset=&config=&split=&where=&orderby=&offset=&length=` - `List parquet shards`: `/parquet?dataset=` - `Get size totals`: `/size?dataset=` - `Get column statistics`: `/statistics?dataset=&config=&split=` - `Get Croissant metadata (if available)`: `/croissant?dataset=` Pagination pattern: ```bash curl "https://datasets-server.huggingface.co/rows?dataset=stanfordnlp/imdb&config=plain_text&split=train&offset=0&length=100" curl "https://datasets-server.huggingface.co/rows?dataset=stanfordnlp/imdb&config=plain_text&split=train&offset=100&length=100" ``` When pagination is partial, use response fields such as `num_rows_total`, `num_rows_per_page`, and `partial` to drive continuation logic. Search/filter notes: - `/search` matches string columns (full-text style behavior is internal to the API). - `/filter` requires predicate syntax in `where` and optional sort in `orderby`. - Keep filtering and searches read-only and side-effect free. For CLI-based parquet URL discovery or SQL, use the `hf-cli` skill with `hf datasets parquet` and `hf datasets sql`. ## Creating and Uploading Datasets Use one of these flows depending on dependency constraints. Zero local dependencies (Hub UI): - Create dataset repo in browser: `https://huggingface.co/new-dataset` - Upload parquet files in the repo "Files and versions" page. - Verify shards appear in Dataset Viewer: ```bash curl -s "https://datasets-server.huggingface.co/parquet?dataset=/" ``` Low dependency CLI

Details

Category Coding → generation
Sourcehuggingface/skills
SKILL.mdView on GitHub →
Repo Stars★ 10.7K
Est. per Skill357 (shared across 30 skills from this repo)
DifficultyIntermediate
Risk LevelN/A

Related Skills

Works Well With

Skills from the same repository — often designed to work together