πŸ“„ To deliver as a Word document: open this page in a browser β†’ press Ctrl+P / Cmd+P β†’ choose β€œSave as PDF”, OR open the HTML file directly in Microsoft Word and use File β–Έ Save As β–Έ Word Document (.docx). Both preserve every heading, table and code block.
Technical Documentation Β· Delivered to Client

HuggingFace Hub
Startup Intelligence Pipeline

Local / VS Code edition β€” end-to-end data collection, checkpointing and reporting for AI-startup activity on the Hugging Face Hub

Document titleHuggingFace Hub Startup Intelligence Pipeline β€” Technical & Operational Documentation
Document version1.0
Date issued28 September 2026
Prepared forClient Data / Research Team
Prepared byData Engineering
Applies tohf_intelligence_pipeline.py (LOCAL / VS Code edition, no cloud storage)
AudienceData engineers, analysts and operators who will run, schedule and maintain the pipeline
StatusFinal β€” ready for handover and execution

1. Purpose and Summary

The HuggingFace Hub Startup Intelligence Pipeline is a single-file Python program that automatically collects, checkpoints and exports structured intelligence about AI startup and organisation activity on the Hugging Face Hub (HF). It runs entirely on a local machine or workstation β€” there is no dependency on Google Cloud Storage, GCS credentials or any cloud SDK.

On each run the pipeline gathers:

  1. Every benchmark leaderboard on the Hub, with a full metadata dump for each ranked model.
  2. Every model created Hub-wide within the run's date window.
  3. Every dataset created Hub-wide within the run's date window.
  4. The Hub's daily research-paper feed for every day in the window.
  5. The complete historical footprint (all models and datasets ever published) of every organisation discovered in items 2 and 3.

Data is progressively written to JSONL checkpoint files (one JSON object per line) so that a long crawl can be stopped at any moment and resumed without re-fetching work already done. At the end of a run, the pipeline consolidates everything into flat CSV tables and a single Excel workbook for analysts.

Why JSONL rather than CSV during collection Records from the Hub do not all share the same fields β€” one model's metadata can contain keys another's does not. CSV assumes a fixed column count per file and silently corrupts data when that assumption breaks. JSONL makes no such assumption, so nothing is lost during the crawl. The conversion to fixed columns happens only at the end, where pandas safely fills blanks for whatever a record did not have.

1.1 Key design principles

PrincipleWhat it means in practice
Dynamic discoveryNo org names, base models, benchmark names or company lists are hardcoded. Everything is discovered from live Hub activity in the date window.
Restartable by designEvery stage appends to its own checkpoint and remembers which items it has already handled. Ctrl+C is always safe.
Resume cursorsFor the two large crawls (models, datasets), the exact next-page URL is saved so a restart jumps straight back to that page instead of re-walking thousands of pages already seen.
Automatic date windowNo manual date editing. The run window is computed from the previous successful run (see Section 9).
Rate-limit toleranceExponential backoff on transient errors and unlimited patient waiting on HTTP 429 responses.
Full-field captureWhere metadata is available it is stored in full, not cherry-picked into a few columns.

2. What the Pipeline Collects

The pipeline is organised into five numbered collection stages followed by an export stage. Each stage is independent: if one fails or is interrupted, the others keep their own checkpoints.

# Stage What is fetched Scope and behaviour
1 Benchmark leaderboards Every dataset on the Hub flagged as a benchmark, and every ranked entry on each benchmark's leaderboard. Each entry is enriched with a full model_info dump (including the model's declared base model). Global β€” not restricted by the run's date window. Each benchmark is fetched once and then permanently marked as done.
2 All new models, Hub-wide Every model repository created between the effective start date and now, with a full metadata dump and any declared base model from its card. Hub-wide. Paginated via REST, sorted newest-first, with a saved page cursor and a "full coverage" marker.
3 All new datasets, Hub-wide Every dataset repository created in the date window, full metadata dump. Identical mechanics to stage 2, with its own independent checkpoints.
4 Daily papers The Hub's curated daily paper feed for every single calendar day in the window β€” title, summary, authors, identifiers, upvotes and paper metadata. One API call per day, per date. Days are remembered so a re-run only fetches missing days.
5 Organisation footprint The entire historical model and dataset catalogue of every organisation discovered in stages 2 and 3 β€” i.e. not just what is new, but everything that organisation has ever published. Org list is derived dynamically from observed namespaces. One pass per org; orgs are recorded as completed as they finish.
β€” Export Consolidation of all checkpoints into CSV files plus one Excel workbook containing one sheet per source. Runs at the end of every execution, whether the crawl completed or was interrupted.
Practical note on scale Stage 5 is by far the largest stage. Because organisation names are discovered dynamically rather than from a fixed list, a single run can surface tens of thousands of namespaces, and each namespace requires at least one and usually two paginated API calls. Budget accordingly β€” see Section 13.

3. How It Works β€” Architecture and Data Flow

3.1 Execution flow

main()
  β”‚
  β”œβ”€ Resolve START_DATE   ← pipeline_state.json (or 2024-06-01 anchor on first run)
  β”‚   END_DATE = now (UTC)
  β”‚
  β”œβ”€ Record row-count snapshot of every master JSONL   (used later to compute deltas)
  β”‚
  β”œβ”€ Stage 1  fetch_all_benchmark_leaderboards()   β†’ 1_leaderboards.jsonl
  β”œβ”€ Stage 2  fetch_all_new_models()               β†’ 2_new_models.jsonl
  β”œβ”€ Stage 3  fetch_all_new_datasets()             β†’ 3_new_datasets.jsonl
  β”œβ”€ Stage 4  fetch_all_daily_papers()             β†’ 4_daily_papers.jsonl
  β”œβ”€ Stage 5  fetch_org_profiles()                 β†’ 5_org_models.jsonl
  β”‚                                                  5_org_datasets.jsonl
  β”‚
  β”œβ”€ Compute this run's delta rows for every source β†’ *_new_<date>.jsonl (temporary)
  β”œβ”€ export_all()        β†’ master CSV copies + HuggingFace_Intelligence_Export.xlsx
  β”œβ”€ export_range_only() β†’ HuggingFace_Intelligence_Export_<start>_to_<end>.xlsx
  β”œβ”€ Delete temporary delta JSONL files
  └─ If the whole run completed: save pipeline_state.json   (advances the next run's start date)

3.2 Checkpointing model

Checkpoints are append-only JSONL files held open for the duration of the run. Writes are flushed after every batch, so partial data is never lost. Three distinct kinds of remembered state exist:

State typeStored asPurpose
Done-keysDerived by re-reading the checkpoint file itself"Which items have I already saved?" β€” the id/benchmark/date field of each stored record is used as the key. No extra bookkeeping file needed.
Page resume cursor*_resume_cursor.json (stages 2 and 3 only)The exact next-page URL from the Hub's Link header, so a restart continues mid-crawl rather than from the top.
Full-coverage marker*_full_coverage_state.json (stages 2 and 3 only)"I have already crawled all the way back to date X once." Future runs then only scan back to X instead of the entire history.
Why the resume cursor matters Walking from the newest model back to 2024 can mean hundreds of thousands of paginated requests. Without a saved cursor, every restart would re-walk every page already visited just to re-confirm records it already holds. With the cursor, a restart lands on the exact page it left off on.

4. System Requirements

4.1 Hardware and operating system

ItemRequirement
Operating systemWindows 10/11, macOS 12+, or a modern Linux distribution (Ubuntu 20.04+ tested pattern)
PythonPython 3.9 or newer (3.10/3.11 recommended)
Disk spacePlan for at least 20 GB free. The stage-5 org footprint files dominate disk use and grow with each new organisation processed. A 50 GB margin is recommended for unattended operation.
RAM8 GB minimum, 16 GB recommended. See the memory note in Section 15.
NetworkContinuous outbound HTTPS access to huggingface.co. Corporate proxies/firewalls must allow this host.
RuntimeA first-time full crawl is a long-running job measured in days, not minutes if the date window is wide and many organisations are discovered. Incremental runs are far shorter.

4.2 Software dependencies

Four third-party packages are required. requests, json, os, time, shutil and datetime that the script also imports are part of the Python standard library.

PackageMinimum versionUsed for
pandas2.1.0Converting JSONL checkpoints into DataFrames, writing CSV and Excel. Version 2.1+ is required because the script uses DataFrame.map for Excel-safe cell cleaning.
requests2.31.0All direct REST calls to the Hub API, including pagination and HTTP 429 handling.
huggingface_hubLatest availableHigh-level helpers: HfApi.model_info, list_datasets(benchmark=True) and get_dataset_leaderboard. Must be a version in which get_dataset_leaderboard exists (see Appendix B for a check).
openpyxl3.1.2The Excel writer engine and the illegal-character sanitiser used before writing text into Excel cells.

4.3 Access requirements

5. Installation and First-Time Setup

Follow these steps once on the machine that will run the pipeline. No prior Hugging Face tooling is assumed.

Step 1 β€” Install Python

Install Python 3.10 or 3.11 from python.org (Windows/macOS) or your distribution's package manager (Linux). Confirm it is available:

python --version
# expected: Python 3.10.x  (or 3.11.x / 3.12.x)

On some Linux systems the command is python3. Use whichever is present throughout.

Step 2 β€” Create the project folder and a virtual environment

Place hf_intelligence_pipeline.py in a dedicated folder, then create an isolated environment so the pipeline's dependencies never conflict with anything else on the machine.

mkdir hf-intelligence
cd hf-intelligence
# copy hf_intelligence_pipeline.py into this folder

python -m venv .venv

# Windows (Command Prompt)
.venv\Scripts\activate

# Windows (PowerShell)
.venv\Scripts\Activate.ps1

# macOS / Linux
source .venv/bin/activate

Step 3 β€” Install dependencies

Create a file named requirements.txt in the same folder with exactly this content:

pandas>=2.1.0
requests>=2.31.0
huggingface_hub>=0.30.0
openpyxl>=3.1.2

Then install:

pip install -r requirements.txt
If get_dataset_leaderboard is missing Run pip install -U huggingface_hub to pull the newest release, then re-run the pre-flight check in Appendix B.

Step 4 β€” Create and install the Hugging Face access token

  1. Sign in at huggingface.co.
  2. Open huggingface.co/settings/tokens β†’ Create new token.
  3. Choose Read permission. Name it something traceable, e.g. startup-intelligence-pipeline.
  4. Copy the token (it is shown only once).
Do not paste the token into the source file The delivered version of the script contains a token literal in the HF_TOKEN line. Storing a live credential inside source code is unsafe: it leaks the moment the file is copied, emailed, committed to Git or shared for review. Rotate that token immediately (revoke it in Settings β–Έ Access Tokens and issue a new one) and supply the replacement through an environment variable instead. Full instructions are in Section 14.

Step 5 β€” Supply the token safely (recommended)

Replace the HF_TOKEN = "hf_..." line near the top of the script with:

HF_TOKEN = os.environ.get("HF_TOKEN", "").strip()

if not HF_TOKEN:
    raise SystemExit(
        "ERROR: the HF_TOKEN environment variable is not set. "
        "Set it before running the pipeline (see documentation, Section 14)."
    )

Then set the variable in the environment that will run the pipeline:

# macOS / Linux (current shell session)
export HF_TOKEN="hf_your_new_read_token"

# Windows PowerShell (current session)
$env:HF_TOKEN = "hf_your_new_read_token"

# Windows Command Prompt (current session)
set HF_TOKEN=hf_your_new_read_token

For a persistent setting use your shell profile (Linux/macOS) or setx / System Properties β–Έ Environment Variables (Windows). Never commit it to version control.

Step 6 β€” Pre-flight check

Before the first real run, execute the small verification script in Appendix B. It confirms the token works, the Hub is reachable, the date window has been resolved correctly and all required library functions are available.

Step 7 β€” First run

Leave FRESH_START = False and run the pipeline exactly as described in Section 7. Monitor the console output; the first run is expected to be long.

6. Configuration Reference

All configuration lives in a single block near the top of hf_intelligence_pipeline.py, under the comment # CONFIG. Nothing below the configuration block needs to be edited for normal operation.

Constant Default Description and guidance
HF_TOKEN placeholder Must be set. The Hugging Face read token. Determines both authentication and the request rate-limit tier. See Section 14 for the secure way to provide it.
OUTPUT_DIR HuggingFace_Intelligence_Output Folder, relative to the working directory, that receives every checkpoint, CSV and workbook. Created automatically if missing. Change it to relocate all output (e.g. to a large data drive such as D:\hf_data).
_ORIGINAL_ANCHOR_START_DATE 2024-06-01 (UTC) The historical backstop used only on a genuinely first run, or after a FRESH_START wipe. Change it only if the client wants the initial backfill to begin at a different date.
START_DATE resolved at runtime Do not edit. Populated at start-up from pipeline_state.json, falling back to the anchor date.
END_DATE datetime.now(UTC) Do not edit. Always "the moment the run started". This is what removes all manual date handling.
FRESH_START False When True, deletes the entire output folder β€” every checkpoint, cursor, state file and export β€” before starting. See the warning below.
REQUEST_DELAY 0.2s with token
1.5s without
Deliberate pause between API calls. Lowered automatically when a token is present. Increase it (e.g. to 0.5) if the run is repeatedly hitting HTTP 429 responses.
MAX_RETRIES 4 Attempts allowed for genuine (non-rate-limit) request failures before the page or item is abandoned. Rate-limit waiting is not capped by this value.
PAGE_LIMIT 100 Records requested per API page for stages 2, 3 and 5. 100 is the Hub's standard practical page size; increasing it is not recommended.
RATE_LIMIT_BACKOFF_CAP 60 seconds Upper bound on a single rate-limit wait. The script will keep retrying indefinitely; this only caps how long each individual pause lasts.
FRESH_START destroys data β€” read before use Setting FRESH_START = True performs an unconditional shutil.rmtree() on the whole output folder. That removes the master checkpoints, the resume cursors, the coverage markers and pipeline_state.json. Consequences:
  • All previously collected history is deleted and must be re-crawled from the anchor date.
  • Because pipeline_state.json is gone, the next run reverts to the 2024-06-01 anchor, producing a very large backfill.
  • There is no undo. Copy the output folder to a backup location first.
Standard practice: use FRESH_START = True only when the client deliberately wants a clean crawl with a changed window. Set it back to False immediately afterwards.
Changing the date range without a full wipe In normal operation you do not change dates at all: bumping END_DATE is unnecessary because it is always "now". To pull a new period of data, simply re-run the script. Only a deliberate re-scope of history requires FRESH_START.

7. Running the Pipeline

7.1 Standard run

With the virtual environment active and the token set:

cd hf-intelligence
# activate the venv (see Section 5, Step 2) if not already active

python hf_intelligence_pipeline.py

Keep the terminal open. The script prints a start-up banner showing the resolved date window, then progress lines for each stage.

7.2 What the start-up banner tells you

[Pipeline State] Last successful run finished through 2026-09-21 β€” this run will pick up from there.

Configured range for this run: 2026-09-21 -> 2026-09-28 (END_DATE = now, auto)
  Models will actually scan:   2026-09-21 -> 2026-09-28
  Datasets will actually scan: 2026-09-21 -> 2026-09-28
REQUEST_DELAY in effect: 0.2s between calls.

Read this carefully on every run. It is the authoritative statement of what will actually be fetched.

7.3 Stopping and resuming

Press Ctrl+C at any time. The script catches the interruption, closes all open checkpoint files cleanly and reports that checkpoints are safe. Simply run the same command again to resume:

Interrupted runs do not advance the date window pipeline_state.json is updated only when all five stages complete without interruption. An interrupted run therefore never accidentally skips data on the next attempt.

7.4 Incremental / weekly run

No configuration change is needed. After a fully successful run, the next execution automatically starts from the previous run's end date and ends at the current moment. This is the intended steady-state operating mode.

7.5 Fresh crawl

  1. Back up the output folder.
  2. Set FRESH_START = True.
  3. Run the script once.
  4. Set FRESH_START = False again before any subsequent run.

7.6 Expected console output (abridged)

=== [1] Benchmark leaderboards (ALL benchmarks, ALL entries) ===
Found 214 total benchmark datasets on the Hub.
  [1/214] SWE-bench/SWE-bench_Verified
    -> +38 entries saved for 'SWE-bench/SWE-bench_Verified' (running total this run: 38)

=== [2] ALL new models Hub-wide (2026-09-21 -> 2026-09-28) ===
  [day 8/8] Scanning 2026-09-28...
  [day 7/8] 2026-09-28 complete β€” 1,412 new models found that day (1,412 total models so far).
  Reached models older than 2026-09-21 after scanning 11,309 records. Stopping crawl.
  [Resume Cursor] Crawl reached effective start date β€” cursor cleared.
  [Full Coverage] Saved β€” next run will only scan back to 2026-09-28 (1-day overlap buffer).

=== [4] Daily papers β€” every day 2026-09-21 -> 2026-09-28 ===
  [day 1/8] 2026-09-21 β€” 34 papers (running total this run: 34)

=== [5] Org footprint β€” discovered dynamically, full historical models+datasets ===
Discovered 6,204 distinct orgs/namespaces from recent Hub activity.
  [200/6204] ... 14,882 models, 3,105 datasets saved this run so far.

=== Combining JSONL checkpoints into CSV + Excel ===
Finished β€” full run completed successfully.

8. Output Files Reference

Everything is written into OUTPUT_DIR (default HuggingFace_Intelligence_Output/). Files fall into four groups: master data, control/state, per-run exports and temporary files.

8.1 Master data checkpoints (append-only, full history)

FileTypeContents
1_leaderboards.jsonl (+ .csv)DataOne row per leaderboard entry across every benchmark, with the entry's own fields plus the full model metadata prefixed model_.
2_new_models.jsonl (+ .csv)DataEvery model repository discovered Hub-wide in the window, with full metadata and a resolved base_model field.
3_new_datasets.jsonl (+ .csv)DataEvery dataset repository discovered Hub-wide in the window.
4_daily_papers.jsonl (+ .csv)DataDaily papers, one row per paper, each stamped with its Date. Days with no papers receive a single placeholder row containing only the date, so the day is recorded as handled.
5_org_models.jsonl (+ .csv)DataThe complete historical model catalogue of every discovered organisation.
5_org_datasets.jsonl (+ .csv)DataThe complete historical dataset catalogue of every discovered organisation.

8.2 Control and state files (do not edit by hand)

FilePurpose
pipeline_state.jsonThe single source of truth for "when did this last finish successfully". Holds last_successful_run_end_date and a saved_at timestamp. Written only after a complete run. Deleting it resets the pipeline to the 2024-06-01 anchor.
2_new_models_resume_cursor.json
3_new_datasets_resume_cursor.json
Saved next-page URL plus the progress date reached, so a restart continues mid-crawl. Automatically deleted once a crawl reaches its start date.
2_new_models_full_coverage_state.json
3_new_datasets_full_coverage_state.json
Records the date through which a complete crawl has been confirmed. Future runs scan back only to that date rather than the full history.
1_leaderboards_skipped.txtBenchmarks that returned no leaderboard or could not be parsed. They are remembered so the script does not retry them each run. Useful as a data-quality audit list.
5_orgs_completed.txtOne organisation name per line, appended as each org finishes. This is what makes stage 5 restartable.

8.3 Per-run exports

FilePurpose
HuggingFace_Intelligence_Export.xlsxThe master workbook: rebuilt at the end of every run from the full cumulative history. Six sheets β€” 1. Leaderboards, 2. New Models, 3. New Datasets, 4. Daily Papers, 5. Org Models, 6. Org Datasets. Always the complete dataset to date.
HuggingFace_Intelligence_Export_<start>_to_<end>.xlsxThe range-only workbook: contains only the rows added during this run, with the actual date window encoded in the filename. This is the file to hand to analysts asking "what is new since last time".
*.csvA CSV copy of every sheet, written alongside the workbook. Always use the CSV for analysis or bulk loading β€” see the Excel limits below.

8.4 Temporary files

Files named <source>_new_<YYYY-MM-DD>.jsonl are the per-run delta rows. They exist only long enough to feed the range-only export and are then deleted automatically. Their data is preserved in two places: the permanent master checkpoint and the range-only workbook. If a run is interrupted before the cleanup step, these files may remain on disk; they are safe to delete manually.

Excel row and width limits Excel supports a maximum of 1,048,576 rows per sheet. The script caps Excel output at 1,000,000 rows per sheet and prints a truncation warning when a source exceeds that. The CSV copy and the JSONL checkpoint always contain the complete data β€” no rows are ever lost from those. Some sheet may still be very wide (many hundreds of columns) because fields vary between records; consider loading the CSVs directly into a database or BI tool for serious analysis.

9. Understanding the Date and Resume Logic

This is the part of the pipeline most likely to cause confusion, so it is documented explicitly. Three separate mechanisms interact.

9.1 Mechanism 1 β€” the run window (pipeline_state.json)

SituationResulting window
No pipeline_state.json (very first run)2024-06-01 β†’ now. A large initial backfill.
pipeline_state.json presentLast successful run's end date β†’ now.
Previous run was interruptedState file untouched, so the window is unchanged from the last success β€” no data is skipped.

9.2 Mechanism 2 β€” full-coverage markers (stages 2 and 3 only)

Stages 2 and 3 keep their own *_full_coverage_state.json. Once a crawl has run all the way back to its start date, the script records the current end date there. On subsequent runs those stages scan back only to that recorded date β€” giving a deliberate one-day overlap buffer to catch late-registering repositories β€” instead of re-walking the entire history from 2024.

The run window and the scan window can differ The printed banner shows both. A run may have a window of 2026-09-21 β†’ 2026-09-28 at the pipeline level while stages 2 and 3 only scan 2026-09-27 β†’ 2026-09-28, because everything earlier was already covered. This is correct and intended behaviour, not an error.

9.3 Mechanism 3 β€” resume cursors (stages 2 and 3 only)

While a crawl is in progress, the next-page URL from the Hub's Link response header is written to *_resume_cursor.json after every page. On restart, the crawl resumes directly at that URL. The cursor is deleted when the crawl reaches its start date, because at that point it is no longer needed.

9.4 Date handling and time zones

10. Data Fields and Schema

The pipeline deliberately stores full metadata rather than a fixed set of columns. The tables below list the fields that can be relied upon, plus the enrichment columns the pipeline itself adds.

10.1 Pipeline-added columns (present in every relevant source)

ColumnAppears inMeaning
Startup NameSources 1, 2, 3, 5The repository owner namespace β€” for a model acme/llama-tune this is acme. In source 1 it is derived from the ranked model's ID; in source 5 it is the organisation being profiled. This is the primary join key for company-level analysis.
base_modelSources 2 and 5 (models)The base-model lineage declared in the model's own card metadata. None when the card does not declare one. May be a list when a card declares several.
model_base_modelSource 1The same lineage information, sourced from the enriched model_info dump for the ranked model.
BenchmarkSource 1The benchmark dataset ID the entry belongs to.
DateSource 4The calendar day the daily-paper row belongs to (YYYY-MM-DD).
model_*Source 1Every field from the enriched model_info call is prefixed model_ so ranked-model metadata cannot collide with leaderboard-entry fields.

10.2 Leaderboard entry fields (source 1)

These come straight from the Hub's leaderboard object for each ranked model:

FieldDescription
rankPosition on the leaderboard.
model_idFull model identifier, e.g. org/model-name.
valueThe benchmark score for that entry.
verifiedWhether the result has been independently verified.
authorUser or organisation object that submitted or owns the result.
sourceWhere the result came from (model card, external submission, etc.).
filenamePath to the evaluation-results YAML file inside the benchmark repository.
pull_requestPull-request number for the submission, when applicable.
notesOptional free-text notes attached to the entry.

Reference: Hugging Face β€” Accessing Benchmark Leaderboard Data.

10.3 Model fields (sources 2 and 5)

Sourced from the Hub's model listing in full-metadata mode. Commonly present: id, createdAt, lastModified, likes, downloads, tags, pipeline_tag, library_name, private, gated, siblings (file list), cardData (parsed model card: licence, datasets, metrics, language, base model), and config. Fields vary from model to model, which is precisely why JSONL is used during collection.

10.4 Dataset fields (sources 3 and 5)

Commonly present: id, createdAt, lastModified, likes, downloads, tags, private, gated, description, siblings, cardData.

10.5 Daily paper fields (source 4)

Each row carries the added Date column plus the raw paper object, typically including paper (with id, title, summary, authors, publishedAt, upvotes, external identifiers), title, thumbnail, numComments and submittedOnDailyAt.

Nested values Values that are themselves objects or lists (authors, file lists, card metadata) are stored as nested JSON inside the cell rather than being flattened into separate columns. Analysts working in a notebook can parse them with json.loads; the pipeline intentionally does not guess a flattening scheme.

11. API Usage and Rate Limits

11.1 Endpoints contacted

See Appendix A for the full list. All traffic is outbound HTTPS to huggingface.co and all requests carry the bearer token.

11.2 Why a token is effectively mandatory

Hugging Face applies rate limits to all Hub API calls, counted over fixed five-minute windows. Published tiers (as of the vendor's rate-limits documentation) are:

PlanAPI requests / 5-minute window
Anonymous (per IP address)500 (subject to change)
Free user (signed in with a token)1,000 (subject to change)
PRO user2,500
Team organisation3,000
Enterprise organisation6,000
Enterprise Plus organisation10,000 (up to 100,000 with higher limits enabled)

Source: Hugging Face β€” Hub Rate limits. Figures are the vendor's published values at the time of writing and may change; verify before capacity planning.

This pipeline makes a very large number of calls. Without a token it is limited to the anonymous tier and the script deliberately raises its inter-request delay to 1.5 seconds, which makes a full crawl impractically slow. For a client running this at scale, a Team or Enterprise organisation token is strongly recommended β€” it raises the ceiling far more than any pacing change can.

11.3 How the pipeline handles throttling

SignalPipeline behaviour
HTTP 429 (rate limited)Waits β€” honouring the server's Retry-After header when present, otherwise backing off progressively up to RATE_LIMIT_BACKOFF_CAP (60s) β€” and retries without limit. Being throttled pauses the run; it does not abort it.
Transient network/HTTP errorsRetried up to MAX_RETRIES (4) times with doubling backoff, then the current page or item is abandoned and the run continues.
Any other non-200 statusLogged clearly and the current crawl stops cleanly, leaving all checkpoints intact for the next run.
Steady-state pacingREQUEST_DELAY between calls β€” 0.2s with a token, 1.5s without.
Sizing expectation Total request volume is roughly: (benchmarks) + (pages to walk in stages 2 and 3) + (days in window) + 2 Γ— (number of organisations discovered). Stage 5 alone can dominate. At a 0.2s delay the floor is about 18,000 requests per hour of wall-clock running time, before allowing for throttling pauses and page latency β€” treat any estimate as indicative and measure the first real run.

12. Troubleshooting Guide

SymptomLikely causeResolution
401 Unauthorized / 403 Forbidden in the log Token missing, mistyped, expired or revoked Confirm the HF_TOKEN environment variable is visible to the process (echo $HF_TOKEN / echo %HF_TOKEN%). Issue a fresh read token and retry.
Frequent HTTP 429 lines, run appears stalled Hitting the account's 5-minute quota This is normal and self-correcting β€” the script waits and continues. To reduce frequency, raise REQUEST_DELAY to 0.5, or move to a higher-tier organisation token.
get_dataset_leaderboard is missing / AttributeError Outdated huggingface_hub pip install -U huggingface_hub and re-run the Appendix B check.
"No orgs discovered yet β€” run steps 2/3 first" Stage 5 executed with no model/dataset checkpoints present (typical after a FRESH_START that was interrupted early) Let stages 2 and 3 complete, then re-run. Expected behaviour, not an error.
"No benchmarks found (or the call failed after retries)" Transient API failure, or the token cannot read the benchmark listings Re-run. If it persists, verify network access to huggingface.co and that the token is valid.
Some benchmarks listed in 1_leaderboards_skipped.txt Those datasets have no leaderboard, or it could not be parsed Expected and harmless β€” benchmarks without a leaderboard simply have no data. Review the file as a data-quality note; delete a line to force a retry.
Excel truncation warning printed A source exceeded the 1,000,000-row Excel cap None needed. Use the CSV copy, which is complete. For a full dataset of that size, load the CSV or JSONL into a database instead of Excel.
openpyxl / IllegalCharacterError Control characters in Hub-provided text Already handled β€” the exporter sanitises illegal characters before writing. If it recurs, confirm openpyxl>=3.1.2 is installed.
AttributeError: 'DataFrame' object has no attribute 'map' pandas older than 2.1 pip install -U "pandas>=2.1.0".
Workbook is very slow to open / huge file size Hundreds of columns and many rows Expected for this dataset shape. Open the CSVs in a notebook or BI tool instead of Excel for day-to-day analysis.
Machine slows down or the process is killed on a long run Accumulating in-memory "already done" key sets, or OS memory pressure Stop and re-run β€” it will resume. Avoid running other heavy workloads concurrently; see Section 15 for the memory note.
Run finished but no new rows appeared Window is genuinely empty, or the stage was already at full coverage Check the banner's scan windows. The range-only workbook will contain a "No New Data" placeholder sheet explaining that no rows were fetched in the period.
Unexpectedly huge backfill after a code change FRESH_START left at True, or pipeline_state.json deleted Verify FRESH_START = False. If the state file was lost, the next clean run re-establishes it.

13. Operational Runbook (Scheduling)

13.1 Recommended cadence

CadenceRationale
WeeklyThe natural rhythm for this dataset. Each run reports exactly what appeared since the last successful run, and the range-only workbook filename states that period explicitly β€” ideal for a Monday-morning briefing pack.
MonthlyAcceptable if the client only needs periodic snapshots. Runs take longer because the window is wider, and the chance of hitting rate limits is higher.
More often than dailyNot advisable. The Hub's creation volume in a sub-day window is small, and repeated runs consume quota without adding much data.

13.2 First run versus steady state

13.3 Unattended execution (Linux / macOS β€” cron)

# Run every Monday at 03:00, logging to a dated file
0 3 * * 1 cd /opt/hf-intelligence && \
  HF_TOKEN="hf_..." /opt/hf-intelligence/.venv/bin/python \
  hf_intelligence_pipeline.py >> logs/run_$(date +\%F).log 2>&1

Prefer putting the token in a root-owned environment file that cron sources, rather than inline as shown, so it does not appear in the crontab or process list.

13.4 Unattended execution (Windows β€” Task Scheduler)

  1. Create a new task with a weekly trigger.
  2. Action: Start a program β†’ program C:\hf-intelligence\.venv\Scripts\python.exe; arguments hf_intelligence_pipeline.py; start-in C:\hf-intelligence.
  3. Set the HF token as a system or user environment variable (System Properties β–Έ Advanced β–Έ Environment Variables) so the task inherits it. Do not embed it in the task arguments.
  4. Enable Run task as soon as possible after a scheduled start is missed so an outage does not silently skip a week.
  5. Set a generous "Stop the task if it runs longer than…" value, or clear it entirely for the first full backfill.

13.5 Monitoring and logging

13.6 Backup

Back up the whole output folder on the same cadence as the pipeline β€” the master checkpoints are the accumulated research asset and are not reproducible without re-crawling. Because the crawl is incremental, losing pipeline_state.json and the coverage markers alongside the data means a full 2024-onwards re-crawl.

14. Security and Token Management

Action required before handover β€” token rotation The delivered script contains a live Hugging Face token written directly into source code (HF_TOKEN = "hf_…"). Anyone holding a copy of this file holds that credential. Remediation:
  1. Go to huggingface.co/settings/tokens and revoke the token currently in the file.
  2. Create a replacement read-scoped token.
  3. Supply it via the HF_TOKEN environment variable (Section 5, Step 5).
  4. Confirm no copy of the old token remains in the script, any notes, screenshots, chat threads or version-control history.
Treat the credential in the delivered file as already compromised.

14.1 Token handling rules

RuleReason
Use a read-scoped tokenThe pipeline only reads public data. A read token cannot modify or delete anything even if it leaks.
Never commit the token to GitIt stays recoverable in history forever. If the project is version-controlled, add the venv, output folder and any .env file to .gitignore.
One token per deploymentLets you revoke a single environment without disturbing others, and makes usage traceable in the account's activity log.
Rotate on a schedule (e.g. every 90 days) and on staff changeLimits the blast radius of any undisclosed exposure.
Prefer an organisation token for productionHigher rate limits, and the token belongs to the organisation rather than an individual β€” so it survives staff turnover.
Restrict file-system access to the output folderIt contains a complete profile of third-party organisations. Apply the client's normal data-handling policy.

14.2 Data handling and compliance notes

15. Known Limitations and Recommendations

The following behaviours are by design or inherent to the approach. They are documented so the client can plan around them rather than treat them as defects.

LimitationDetail and recommended mitigation
Leaderboards are snapshotted once per benchmark Once a benchmark has been fetched, it is permanently marked done, so scores added to that benchmark later are not picked up. Mitigation: to refresh a benchmark, remove its dataset ID from 1_leaderboards.jsonl (or delete the whole file to refresh everything) and re-run. Document the intended refresh policy with the client up front.
Organisation footprint is captured once per organisation An organisation already recorded in 5_orgs_completed.txt is not re-profiled, so repositories it publishes later will not appear in source 5 β€” though genuinely new repositories do appear in sources 2 and 3. Mitigation: for a periodic re-profile, delete the chosen org names from 5_orgs_completed.txt (and optionally their rows from the two 5_org_*.jsonl files) before a run.
Stage 5 has no upper bound on scope Organisation discovery is dynamic, so a wide date window can surface tens of thousands of namespaces, each requiring one or two API calls. The stage is designed to be interruptible and resumable; split the work across several sittings if necessary.
In-memory "already done" key sets grow during a run Each stage holds the set of processed IDs in memory while it runs. On very large crawls this consumes a noticeable amount of RAM. Mitigation: 16 GB RAM recommended; stopping and restarting rebuilds the sets from the checkpoint files, which releases memory.
Model-info cache is per-run only Leaderboard enrichment caches model_info responses in memory and does not persist them between runs, so a re-run of stage 1 repeats those lookups. Acceptable given the stage is normally run once per benchmark.
Wide, sparse tables in the export Because records carry different fields, the flat export has many columns and many blanks. Mitigation: prefer the CSVs over Excel, load into a database for repeated querying, and define analysis views rather than working with the raw wide table.
Excel is not a delivery format for full volumes Sheets are capped at 1,000,000 rows by the script (Excel's own limit is 1,048,576). CSV and JSONL always hold the complete data. Agree with the client in advance which format is authoritative for downstream use.
No automated alerting Failures are reported in the console and log file only. Mitigation: add a wrapper script that checks the final log line and sends an email/Teams message on an unexpected exit code.
No schema versioning on the checkpoints Field sets can shift if the Hub changes its API. Existing checkpoints are never rewritten, so a change affects only new rows. Note the date of any Hub API change in the run log.

16. Maintenance and Customisation

16.1 Common changes the client may request

RequirementHow to implement
Move output to another driveChange OUTPUT_DIR to an absolute path, e.g. r"D:\hf_intelligence_data".
Change the historical backfill startChange _ORIGINAL_ANCHOR_START_DATE, then run once with FRESH_START = True and set it back to False.
Re-collect a specific week onlyTemporarily comment out stages 1 and 5 in the __main__ block and set both full-coverage states to the required start date, or delete them β€” see Section 9.
Speed up on a high-limit planLower REQUEST_DELAY to 0.1 and re-run; monitor for 429s.
Add a new data sourceCopy the shape of an existing stage: a dedicated fetch_… function, its own JSONL checkpoint, a done-key field, and an entry in the export's sources list.
Push results to a database or dashboardLoad the exported CSVs (or the JSONL directly) with a downstream job. The checkpoint files are append-only, so a simple incremental load on id works well.

16.2 Upgrade procedure when Hugging Face changes its API

  1. Stop any running job cleanly with Ctrl+C.
  2. Back up the entire output folder.
  3. Update the library: pip install -U huggingface_hub requests pandas openpyxl.
  4. Run the Appendix B pre-flight check β€” it validates endpoint and field access before any long crawl begins.
  5. Run one incremental job and inspect 1_leaderboards.jsonl and 2_new_models.jsonl tails for unexpected empty fields.
  6. Only then resume normal scheduling.

16.3 Support information to capture with any issue report

Appendix A β€” Hugging Face API Endpoints Used

StageEndpoint / methodPurpose
1HfApi.list_datasets(benchmark=True)Enumerate every benchmark dataset on the Hub.
1GET /api/datasets/{dataset_id}/leaderboard
via HfApi.get_dataset_leaderboard()
Fetch the ranked model scores for one benchmark.
1HfApi.model_info(model_id, securityStatus=True)Full metadata dump for each ranked model, including its card data.
2GET /api/models?sort=createdAt&direction=-1&limit=100&full=true&cardData=truePaginated listing of all models, newest first, with full metadata.
3GET /api/datasets?sort=createdAt&direction=-1&limit=100&full=true&cardData=truePaginated listing of all datasets, newest first, with full metadata.
4GET /api/daily_papers?date=YYYY-MM-DDThe curated paper feed for one calendar day.
5GET /api/models?author={org} and GET /api/datasets?author={org}Complete historical catalogue for one organisation.

A.1 Pagination

Stages 2, 3 and 5 follow the Hub's Link response header (rel="next") to walk pages. The next-page URL is saved after every page, which is what makes mid-crawl resumption possible.

A.2 Reference documentation

Appendix B β€” Pre-flight Health Check Script

Save this as preflight_check.py in the project folder and run it with the virtual environment active before the first real run and after any library upgrade. It validates the token, network access, library capabilities and date resolution without collecting any data.

"""
preflight_check.py -- validates the environment for hf_intelligence_pipeline.py
Run:  python preflight_check.py
"""
import os, sys, json
from datetime import datetime, timezone

print("=" * 68)
print("HuggingFace Intelligence Pipeline -- pre-flight check")
print("=" * 68)

ok = True

# 1. Python version ---------------------------------------------------------
major, minor = sys.version_info[:2]
print(f"\n[1] Python version: {major}.{minor}")
if (major, minor) < (3, 9):
    print("    FAIL - Python 3.9 or newer is required.")
    ok = False
else:
    print("    OK")

# 2. Packages ---------------------------------------------------------------
print("\n[2] Required packages")
try:
    import pandas as pd
    print(f"    pandas           {pd.__version__}")
    if pd.__version__ < "2.1":
        print("    FAIL - pandas 2.1+ is required (DataFrame.map).")
        ok = False
except ImportError:
    print("    FAIL - pandas not installed. Run: pip install -r requirements.txt")
    ok = False

try:
    import requests
    print(f"    requests         {requests.__version__}")
except ImportError:
    print("    FAIL - requests not installed.")
    ok = False

try:
    import openpyxl
    print(f"    openpyxl         {openpyxl.__version__}")
except ImportError:
    print("    FAIL - openpyxl not installed.")
    ok = False

try:
    import huggingface_hub
    from huggingface_hub import HfApi
    print(f"    huggingface_hub  {huggingface_hub.__version__}")
except ImportError:
    print("    FAIL - huggingface_hub not installed.")
    ok = False

# 3. Library capability -----------------------------------------------------
print("\n[3] Library capabilities")
api = HfApi()
for name in ("get_dataset_leaderboard", "list_datasets", "model_info"):
    if hasattr(api, name):
        print(f"    OK   HfApi.{name}")
    else:
        print(f"    FAIL HfApi.{name} is missing -- upgrade huggingface_hub.")
        ok = False

# 4. Token ------------------------------------------------------------------
print("\n[4] Authentication")
token = os.environ.get("HF_TOKEN", "").strip()
if not token:
    print("    FAIL - HF_TOKEN environment variable is not set (or is empty).")
    print("           See documentation Section 14 for how to set it safely.")
    ok = False
else:
    print(f"    Token found (length {len(token)}, prefix {token[:6]}...)")
    try:
        who = HfApi(token=token).whoami()
        print(f"    OK   Authenticated as: {who.get('name')} ({who.get('type')})")
    except Exception as e:
        print(f"    FAIL - token rejected by the Hub: {e}")
        ok = False

# 5. Network reachability ---------------------------------------------------
print("\n[5] Network reachability")
try:
    r = requests.get("https://huggingface.co/api/models?limit=1",
                     headers=({"Authorization": f"Bearer {token}"} if token else {}),
                     timeout=15)
    print(f"    OK   GET /api/models returned HTTP {r.status_code}")
except Exception as e:
    print(f"    FAIL - cannot reach huggingface.co: {e}")
    ok = False

# 6. Output directory -------------------------------------------------------
print("\n[6] Output directory")
out_dir = "HuggingFace_Intelligence_Output"
try:
    os.makedirs(out_dir, exist_ok=True)
    probe = os.path.join(out_dir, ".write_test")
    with open(probe, "w") as f:
        f.write("ok")
    os.remove(probe)
    print(f"    OK   '{out_dir}' exists and is writable")
except Exception as e:
    print(f"    FAIL - cannot write to '{out_dir}': {e}")
    ok = False

# 7. Date window ------------------------------------------------------------
print("\n[7] Date window resolution")
state_path = os.path.join(out_dir, "pipeline_state.json")
if os.path.exists(state_path):
    try:
        with open(state_path, encoding="utf-8") as f:
            state = json.load(f)
        print(f"    Previous successful run through: {state.get('last_successful_run_end_date')}")
    except Exception as e:
        print(f"    WARN - could not read {state_path}: {e}")
else:
    print("    No pipeline_state.json yet -> first run will backfill from 2024-06-01.")
print(f"    END_DATE for the next run will be: {datetime.now(timezone.utc).isoformat()}")

# 8. Disk space -------------------------------------------------------------
print("\n[8] Free disk space")
try:
    import shutil
    free_gb = shutil.disk_usage(os.path.abspath(out_dir)).free / (1024 ** 3)
    print(f"    {free_gb:.1f} GB free on the output drive")
    if free_gb < 20:
        print("    WARN - less than 20 GB free; a full crawl needs substantial space.")
except Exception as e:
    print(f"    WARN - could not determine disk space: {e}")

print("\n" + "=" * 68)
print("RESULT:", "ALL CHECKS PASSED - safe to run the pipeline." if ok
      else "ONE OR MORE CHECKS FAILED - fix the items marked FAIL above.")
print("=" * 68)
sys.exit(0 if ok else 1)

Appendix C β€” Glossary

TermDefinition
Anchor dateThe historical backstop (2024-06-01) used for the very first crawl or after a full wipe.
BenchmarkA dataset repo on the Hub that carries a leaderboard ranking models by evaluation score.
base_modelThe upstream model a model was derived from, as declared in its model card metadata.
CheckpointA JSONL file that accumulates collected records so a run can be stopped and resumed safely.
Coverage stateA file recording the date through which a full crawl has already been confirmed, so future runs need not repeat it.
DeltaThe rows added during a single run, as distinct from the cumulative master history.
Done-keysThe set of identifiers already present in a checkpoint, used to skip work on re-runs.
HubThe Hugging Face model, dataset and Space hosting platform.
JSONLJSON Lines β€” one complete JSON object per line. Tolerant of records with differing field sets.
Namespace / orgThe owner portion of a repository ID. Used as the company-level identifier throughout the outputs.
Rate limitThe Hub's cap on API requests, counted over fixed five-minute windows.
Resume cursorThe saved next-page URL that lets a crawl restart mid-way instead of from the top.
Run windowThe START_DATE β†’ END_DATE period a given execution covers.

Appendix D β€” Handover Acceptance Checklist

Use this list to confirm the pipeline is genuinely ready for client operation. Each item should be verified by the client, not assumed.

#ItemVerified by / date
1Token rotated. The token present in the delivered script has been revoked and replaced; the replacement is supplied via environment variable only.
2Python 3.9+ installed and a virtual environment created on the target machine.
3pip install -r requirements.txt completed without error.
4preflight_check.py (Appendix B) reports ALL CHECKS PASSED.
5At least 20 GB free disk space confirmed on the output drive.
6OUTPUT_DIR set to the agreed production location.
7FRESH_START confirmed as False for normal operation.
8First full run executed to completion; pipeline_state.json created.
9Master workbook and range-only workbook both opened successfully and sheet contents spot-checked against the Hub.
10Second (incremental) run executed; banner confirmed it scanned only the new period.
11Scheduling configured (cron / Task Scheduler) with output redirected to a log file.
12Backup routine for the output folder agreed and in place.
13Refresh policy agreed for leaderboards and organisation footprints (Section 15).
14Support/escalation contact and the information-to-capture list (Section 16.3) shared with the client's operations team.
15Client has received this document and confirmed the run instructions were followed unaided at least once.