Local / VS Code edition β end-to-end data collection, checkpointing and reporting for AI-startup activity on the Hugging Face Hub
| Document title | HuggingFace Hub Startup Intelligence Pipeline β Technical & Operational Documentation |
| Document version | 1.0 |
| Date issued | 28 September 2026 |
| Prepared for | Client Data / Research Team |
| Prepared by | Data Engineering |
| Applies to | hf_intelligence_pipeline.py (LOCAL / VS Code edition, no cloud storage) |
| Audience | Data engineers, analysts and operators who will run, schedule and maintain the pipeline |
| Status | Final β ready for handover and execution |
The HuggingFace Hub Startup Intelligence Pipeline is a single-file Python program that automatically collects, checkpoints and exports structured intelligence about AI startup and organisation activity on the Hugging Face Hub (HF). It runs entirely on a local machine or workstation β there is no dependency on Google Cloud Storage, GCS credentials or any cloud SDK.
On each run the pipeline gathers:
Data is progressively written to JSONL checkpoint files (one JSON object per line) so that a long crawl can be stopped at any moment and resumed without re-fetching work already done. At the end of a run, the pipeline consolidates everything into flat CSV tables and a single Excel workbook for analysts.
| Principle | What it means in practice |
|---|---|
| Dynamic discovery | No org names, base models, benchmark names or company lists are hardcoded. Everything is discovered from live Hub activity in the date window. |
| Restartable by design | Every stage appends to its own checkpoint and remembers which items it has already handled. Ctrl+C is always safe. |
| Resume cursors | For the two large crawls (models, datasets), the exact next-page URL is saved so a restart jumps straight back to that page instead of re-walking thousands of pages already seen. |
| Automatic date window | No manual date editing. The run window is computed from the previous successful run (see Section 9). |
| Rate-limit tolerance | Exponential backoff on transient errors and unlimited patient waiting on HTTP 429 responses. |
| Full-field capture | Where metadata is available it is stored in full, not cherry-picked into a few columns. |
The pipeline is organised into five numbered collection stages followed by an export stage. Each stage is independent: if one fails or is interrupted, the others keep their own checkpoints.
| # | Stage | What is fetched | Scope and behaviour |
|---|---|---|---|
| 1 | Benchmark leaderboards | Every dataset on the Hub flagged as a benchmark, and every ranked entry on each benchmark's leaderboard. Each entry is enriched with a full model_info dump (including the model's declared base model). |
Global β not restricted by the run's date window. Each benchmark is fetched once and then permanently marked as done. |
| 2 | All new models, Hub-wide | Every model repository created between the effective start date and now, with a full metadata dump and any declared base model from its card. | Hub-wide. Paginated via REST, sorted newest-first, with a saved page cursor and a "full coverage" marker. |
| 3 | All new datasets, Hub-wide | Every dataset repository created in the date window, full metadata dump. | Identical mechanics to stage 2, with its own independent checkpoints. |
| 4 | Daily papers | The Hub's curated daily paper feed for every single calendar day in the window β title, summary, authors, identifiers, upvotes and paper metadata. | One API call per day, per date. Days are remembered so a re-run only fetches missing days. |
| 5 | Organisation footprint | The entire historical model and dataset catalogue of every organisation discovered in stages 2 and 3 β i.e. not just what is new, but everything that organisation has ever published. | Org list is derived dynamically from observed namespaces. One pass per org; orgs are recorded as completed as they finish. |
| β | Export | Consolidation of all checkpoints into CSV files plus one Excel workbook containing one sheet per source. | Runs at the end of every execution, whether the crawl completed or was interrupted. |
main()
β
ββ Resolve START_DATE β pipeline_state.json (or 2024-06-01 anchor on first run)
β END_DATE = now (UTC)
β
ββ Record row-count snapshot of every master JSONL (used later to compute deltas)
β
ββ Stage 1 fetch_all_benchmark_leaderboards() β 1_leaderboards.jsonl
ββ Stage 2 fetch_all_new_models() β 2_new_models.jsonl
ββ Stage 3 fetch_all_new_datasets() β 3_new_datasets.jsonl
ββ Stage 4 fetch_all_daily_papers() β 4_daily_papers.jsonl
ββ Stage 5 fetch_org_profiles() β 5_org_models.jsonl
β 5_org_datasets.jsonl
β
ββ Compute this run's delta rows for every source β *_new_<date>.jsonl (temporary)
ββ export_all() β master CSV copies + HuggingFace_Intelligence_Export.xlsx
ββ export_range_only() β HuggingFace_Intelligence_Export_<start>_to_<end>.xlsx
ββ Delete temporary delta JSONL files
ββ If the whole run completed: save pipeline_state.json (advances the next run's start date)
Checkpoints are append-only JSONL files held open for the duration of the run. Writes are flushed after every batch, so partial data is never lost. Three distinct kinds of remembered state exist:
| State type | Stored as | Purpose |
|---|---|---|
| Done-keys | Derived by re-reading the checkpoint file itself | "Which items have I already saved?" β the id/benchmark/date field of each stored record is used as the key. No extra bookkeeping file needed. |
| Page resume cursor | *_resume_cursor.json (stages 2 and 3 only) | The exact next-page URL from the Hub's Link header, so a restart continues mid-crawl rather than from the top. |
| Full-coverage marker | *_full_coverage_state.json (stages 2 and 3 only) | "I have already crawled all the way back to date X once." Future runs then only scan back to X instead of the entire history. |
| Item | Requirement |
|---|---|
| Operating system | Windows 10/11, macOS 12+, or a modern Linux distribution (Ubuntu 20.04+ tested pattern) |
| Python | Python 3.9 or newer (3.10/3.11 recommended) |
| Disk space | Plan for at least 20 GB free. The stage-5 org footprint files dominate disk use and grow with each new organisation processed. A 50 GB margin is recommended for unattended operation. |
| RAM | 8 GB minimum, 16 GB recommended. See the memory note in Section 15. |
| Network | Continuous outbound HTTPS access to huggingface.co. Corporate proxies/firewalls must allow this host. |
| Runtime | A first-time full crawl is a long-running job measured in days, not minutes if the date window is wide and many organisations are discovered. Incremental runs are far shorter. |
Four third-party packages are required. requests, json, os, time, shutil and datetime that the script also imports are part of the Python standard library.
| Package | Minimum version | Used for |
|---|---|---|
pandas | 2.1.0 | Converting JSONL checkpoints into DataFrames, writing CSV and Excel. Version 2.1+ is required because the script uses DataFrame.map for Excel-safe cell cleaning. |
requests | 2.31.0 | All direct REST calls to the Hub API, including pagination and HTTP 429 handling. |
huggingface_hub | Latest available | High-level helpers: HfApi.model_info, list_datasets(benchmark=True) and get_dataset_leaderboard. Must be a version in which get_dataset_leaderboard exists (see Appendix B for a check). |
openpyxl | 3.1.2 | The Excel writer engine and the illegal-character sanitiser used before writing text into Excel cells. |
huggingface.co/settings/tokens. A read token is sufficient; no write permission is needed anywhere in this pipeline.Follow these steps once on the machine that will run the pipeline. No prior Hugging Face tooling is assumed.
Install Python 3.10 or 3.11 from python.org (Windows/macOS) or your distribution's package manager (Linux). Confirm it is available:
python --version
# expected: Python 3.10.x (or 3.11.x / 3.12.x)
On some Linux systems the command is python3. Use whichever is present throughout.
Place hf_intelligence_pipeline.py in a dedicated folder, then create an isolated environment so the pipeline's dependencies never conflict with anything else on the machine.
mkdir hf-intelligence
cd hf-intelligence
# copy hf_intelligence_pipeline.py into this folder
python -m venv .venv
# Windows (Command Prompt)
.venv\Scripts\activate
# Windows (PowerShell)
.venv\Scripts\Activate.ps1
# macOS / Linux
source .venv/bin/activate
Create a file named requirements.txt in the same folder with exactly this content:
pandas>=2.1.0
requests>=2.31.0
huggingface_hub>=0.30.0
openpyxl>=3.1.2
Then install:
pip install -r requirements.txt
get_dataset_leaderboard is missing
Run pip install -U huggingface_hub to pull the newest release, then re-run the pre-flight check in Appendix B.
huggingface.co.huggingface.co/settings/tokens β Create new token.startup-intelligence-pipeline.HF_TOKEN line. Storing a live credential inside source code is unsafe: it leaks the moment the file is copied, emailed, committed to Git or shared for review. Rotate that token immediately (revoke it in Settings βΈ Access Tokens and issue a new one) and supply the replacement through an environment variable instead. Full instructions are in Section 14.
Replace the HF_TOKEN = "hf_..." line near the top of the script with:
HF_TOKEN = os.environ.get("HF_TOKEN", "").strip()
if not HF_TOKEN:
raise SystemExit(
"ERROR: the HF_TOKEN environment variable is not set. "
"Set it before running the pipeline (see documentation, Section 14)."
)
Then set the variable in the environment that will run the pipeline:
# macOS / Linux (current shell session)
export HF_TOKEN="hf_your_new_read_token"
# Windows PowerShell (current session)
$env:HF_TOKEN = "hf_your_new_read_token"
# Windows Command Prompt (current session)
set HF_TOKEN=hf_your_new_read_token
For a persistent setting use your shell profile (Linux/macOS) or setx / System Properties βΈ Environment Variables (Windows). Never commit it to version control.
Before the first real run, execute the small verification script in Appendix B. It confirms the token works, the Hub is reachable, the date window has been resolved correctly and all required library functions are available.
Leave FRESH_START = False and run the pipeline exactly as described in Section 7. Monitor the console output; the first run is expected to be long.
All configuration lives in a single block near the top of hf_intelligence_pipeline.py, under the comment # CONFIG. Nothing below the configuration block needs to be edited for normal operation.
| Constant | Default | Description and guidance |
|---|---|---|
HF_TOKEN |
placeholder | Must be set. The Hugging Face read token. Determines both authentication and the request rate-limit tier. See Section 14 for the secure way to provide it. |
OUTPUT_DIR |
HuggingFace_Intelligence_Output |
Folder, relative to the working directory, that receives every checkpoint, CSV and workbook. Created automatically if missing. Change it to relocate all output (e.g. to a large data drive such as D:\hf_data). |
_ORIGINAL_ANCHOR_START_DATE |
2024-06-01 (UTC) |
The historical backstop used only on a genuinely first run, or after a FRESH_START wipe. Change it only if the client wants the initial backfill to begin at a different date. |
START_DATE |
resolved at runtime | Do not edit. Populated at start-up from pipeline_state.json, falling back to the anchor date. |
END_DATE |
datetime.now(UTC) |
Do not edit. Always "the moment the run started". This is what removes all manual date handling. |
FRESH_START |
False |
When True, deletes the entire output folder β every checkpoint, cursor, state file and export β before starting. See the warning below. |
REQUEST_DELAY |
0.2s with token1.5s without |
Deliberate pause between API calls. Lowered automatically when a token is present. Increase it (e.g. to 0.5) if the run is repeatedly hitting HTTP 429 responses. |
MAX_RETRIES |
4 |
Attempts allowed for genuine (non-rate-limit) request failures before the page or item is abandoned. Rate-limit waiting is not capped by this value. |
PAGE_LIMIT |
100 |
Records requested per API page for stages 2, 3 and 5. 100 is the Hub's standard practical page size; increasing it is not recommended. |
RATE_LIMIT_BACKOFF_CAP |
60 seconds |
Upper bound on a single rate-limit wait. The script will keep retrying indefinitely; this only caps how long each individual pause lasts. |
FRESH_START = True performs an unconditional shutil.rmtree() on the whole output folder. That removes the master checkpoints, the resume cursors, the coverage markers and pipeline_state.json. Consequences:
pipeline_state.json is gone, the next run reverts to the 2024-06-01 anchor, producing a very large backfill.FRESH_START = True only when the client deliberately wants a clean crawl with a changed window. Set it back to False immediately afterwards.
END_DATE is unnecessary because it is always "now". To pull a new period of data, simply re-run the script. Only a deliberate re-scope of history requires FRESH_START.
With the virtual environment active and the token set:
cd hf-intelligence
# activate the venv (see Section 5, Step 2) if not already active
python hf_intelligence_pipeline.py
Keep the terminal open. The script prints a start-up banner showing the resolved date window, then progress lines for each stage.
[Pipeline State] Last successful run finished through 2026-09-21 β this run will pick up from there.
Configured range for this run: 2026-09-21 -> 2026-09-28 (END_DATE = now, auto)
Models will actually scan: 2026-09-21 -> 2026-09-28
Datasets will actually scan: 2026-09-21 -> 2026-09-28
REQUEST_DELAY in effect: 0.2s between calls.
Read this carefully on every run. It is the authoritative statement of what will actually be fetched.
Press Ctrl+C at any time. The script catches the interruption, closes all open checkpoint files cleanly and reports that checkpoints are safe. Simply run the same command again to resume:
pipeline_state.json is updated only when all five stages complete without interruption. An interrupted run therefore never accidentally skips data on the next attempt.
No configuration change is needed. After a fully successful run, the next execution automatically starts from the previous run's end date and ends at the current moment. This is the intended steady-state operating mode.
FRESH_START = True.FRESH_START = False again before any subsequent run.=== [1] Benchmark leaderboards (ALL benchmarks, ALL entries) ===
Found 214 total benchmark datasets on the Hub.
[1/214] SWE-bench/SWE-bench_Verified
-> +38 entries saved for 'SWE-bench/SWE-bench_Verified' (running total this run: 38)
=== [2] ALL new models Hub-wide (2026-09-21 -> 2026-09-28) ===
[day 8/8] Scanning 2026-09-28...
[day 7/8] 2026-09-28 complete β 1,412 new models found that day (1,412 total models so far).
Reached models older than 2026-09-21 after scanning 11,309 records. Stopping crawl.
[Resume Cursor] Crawl reached effective start date β cursor cleared.
[Full Coverage] Saved β next run will only scan back to 2026-09-28 (1-day overlap buffer).
=== [4] Daily papers β every day 2026-09-21 -> 2026-09-28 ===
[day 1/8] 2026-09-21 β 34 papers (running total this run: 34)
=== [5] Org footprint β discovered dynamically, full historical models+datasets ===
Discovered 6,204 distinct orgs/namespaces from recent Hub activity.
[200/6204] ... 14,882 models, 3,105 datasets saved this run so far.
=== Combining JSONL checkpoints into CSV + Excel ===
Finished β full run completed successfully.
Everything is written into OUTPUT_DIR (default HuggingFace_Intelligence_Output/). Files fall into four groups: master data, control/state, per-run exports and temporary files.
| File | Type | Contents |
|---|---|---|
1_leaderboards.jsonl (+ .csv) | Data | One row per leaderboard entry across every benchmark, with the entry's own fields plus the full model metadata prefixed model_. |
2_new_models.jsonl (+ .csv) | Data | Every model repository discovered Hub-wide in the window, with full metadata and a resolved base_model field. |
3_new_datasets.jsonl (+ .csv) | Data | Every dataset repository discovered Hub-wide in the window. |
4_daily_papers.jsonl (+ .csv) | Data | Daily papers, one row per paper, each stamped with its Date. Days with no papers receive a single placeholder row containing only the date, so the day is recorded as handled. |
5_org_models.jsonl (+ .csv) | Data | The complete historical model catalogue of every discovered organisation. |
5_org_datasets.jsonl (+ .csv) | Data | The complete historical dataset catalogue of every discovered organisation. |
| File | Purpose |
|---|---|
pipeline_state.json | The single source of truth for "when did this last finish successfully". Holds last_successful_run_end_date and a saved_at timestamp. Written only after a complete run. Deleting it resets the pipeline to the 2024-06-01 anchor. |
2_new_models_resume_cursor.json3_new_datasets_resume_cursor.json | Saved next-page URL plus the progress date reached, so a restart continues mid-crawl. Automatically deleted once a crawl reaches its start date. |
2_new_models_full_coverage_state.json3_new_datasets_full_coverage_state.json | Records the date through which a complete crawl has been confirmed. Future runs scan back only to that date rather than the full history. |
1_leaderboards_skipped.txt | Benchmarks that returned no leaderboard or could not be parsed. They are remembered so the script does not retry them each run. Useful as a data-quality audit list. |
5_orgs_completed.txt | One organisation name per line, appended as each org finishes. This is what makes stage 5 restartable. |
| File | Purpose |
|---|---|
HuggingFace_Intelligence_Export.xlsx | The master workbook: rebuilt at the end of every run from the full cumulative history. Six sheets β 1. Leaderboards, 2. New Models, 3. New Datasets, 4. Daily Papers, 5. Org Models, 6. Org Datasets. Always the complete dataset to date. |
HuggingFace_Intelligence_Export_<start>_to_<end>.xlsx | The range-only workbook: contains only the rows added during this run, with the actual date window encoded in the filename. This is the file to hand to analysts asking "what is new since last time". |
*.csv | A CSV copy of every sheet, written alongside the workbook. Always use the CSV for analysis or bulk loading β see the Excel limits below. |
Files named <source>_new_<YYYY-MM-DD>.jsonl are the per-run delta rows. They exist only long enough to feed the range-only export and are then deleted automatically. Their data is preserved in two places: the permanent master checkpoint and the range-only workbook. If a run is interrupted before the cleanup step, these files may remain on disk; they are safe to delete manually.
This is the part of the pipeline most likely to cause confusion, so it is documented explicitly. Three separate mechanisms interact.
pipeline_state.json)| Situation | Resulting window |
|---|---|
No pipeline_state.json (very first run) | 2024-06-01 β now. A large initial backfill. |
pipeline_state.json present | Last successful run's end date β now. |
| Previous run was interrupted | State file untouched, so the window is unchanged from the last success β no data is skipped. |
Stages 2 and 3 keep their own *_full_coverage_state.json. Once a crawl has run all the way back to its start date, the script records the current end date there. On subsequent runs those stages scan back only to that recorded date β giving a deliberate one-day overlap buffer to catch late-registering repositories β instead of re-walking the entire history from 2024.
While a crawl is in progress, the next-page URL from the Hub's Link response header is written to *_resume_cursor.json after every page. On restart, the crawl resumes directly at that URL. The cursor is deleted when the crawl reaches its start date, because at that point it is no longer needed.
END_DATE is the UTC time at which the run started.END_DATE are skipped; records older than the effective start date stop the crawl.The pipeline deliberately stores full metadata rather than a fixed set of columns. The tables below list the fields that can be relied upon, plus the enrichment columns the pipeline itself adds.
| Column | Appears in | Meaning |
|---|---|---|
Startup Name | Sources 1, 2, 3, 5 | The repository owner namespace β for a model acme/llama-tune this is acme. In source 1 it is derived from the ranked model's ID; in source 5 it is the organisation being profiled. This is the primary join key for company-level analysis. |
base_model | Sources 2 and 5 (models) | The base-model lineage declared in the model's own card metadata. None when the card does not declare one. May be a list when a card declares several. |
model_base_model | Source 1 | The same lineage information, sourced from the enriched model_info dump for the ranked model. |
Benchmark | Source 1 | The benchmark dataset ID the entry belongs to. |
Date | Source 4 | The calendar day the daily-paper row belongs to (YYYY-MM-DD). |
model_* | Source 1 | Every field from the enriched model_info call is prefixed model_ so ranked-model metadata cannot collide with leaderboard-entry fields. |
These come straight from the Hub's leaderboard object for each ranked model:
| Field | Description |
|---|---|
rank | Position on the leaderboard. |
model_id | Full model identifier, e.g. org/model-name. |
value | The benchmark score for that entry. |
verified | Whether the result has been independently verified. |
author | User or organisation object that submitted or owns the result. |
source | Where the result came from (model card, external submission, etc.). |
filename | Path to the evaluation-results YAML file inside the benchmark repository. |
pull_request | Pull-request number for the submission, when applicable. |
notes | Optional free-text notes attached to the entry. |
Reference: Hugging Face β Accessing Benchmark Leaderboard Data.
Sourced from the Hub's model listing in full-metadata mode. Commonly present: id, createdAt, lastModified, likes, downloads, tags, pipeline_tag, library_name, private, gated, siblings (file list), cardData (parsed model card: licence, datasets, metrics, language, base model), and config. Fields vary from model to model, which is precisely why JSONL is used during collection.
Commonly present: id, createdAt, lastModified, likes, downloads, tags, private, gated, description, siblings, cardData.
Each row carries the added Date column plus the raw paper object, typically including paper (with id, title, summary, authors, publishedAt, upvotes, external identifiers), title, thumbnail, numComments and submittedOnDailyAt.
json.loads; the pipeline intentionally does not guess a flattening scheme.
See Appendix A for the full list. All traffic is outbound HTTPS to huggingface.co and all requests carry the bearer token.
Hugging Face applies rate limits to all Hub API calls, counted over fixed five-minute windows. Published tiers (as of the vendor's rate-limits documentation) are:
| Plan | API requests / 5-minute window |
|---|---|
| Anonymous (per IP address) | 500 (subject to change) |
| Free user (signed in with a token) | 1,000 (subject to change) |
| PRO user | 2,500 |
| Team organisation | 3,000 |
| Enterprise organisation | 6,000 |
| Enterprise Plus organisation | 10,000 (up to 100,000 with higher limits enabled) |
Source: Hugging Face β Hub Rate limits. Figures are the vendor's published values at the time of writing and may change; verify before capacity planning.
This pipeline makes a very large number of calls. Without a token it is limited to the anonymous tier and the script deliberately raises its inter-request delay to 1.5 seconds, which makes a full crawl impractically slow. For a client running this at scale, a Team or Enterprise organisation token is strongly recommended β it raises the ceiling far more than any pacing change can.
| Signal | Pipeline behaviour |
|---|---|
| HTTP 429 (rate limited) | Waits β honouring the server's Retry-After header when present, otherwise backing off progressively up to RATE_LIMIT_BACKOFF_CAP (60s) β and retries without limit. Being throttled pauses the run; it does not abort it. |
| Transient network/HTTP errors | Retried up to MAX_RETRIES (4) times with doubling backoff, then the current page or item is abandoned and the run continues. |
| Any other non-200 status | Logged clearly and the current crawl stops cleanly, leaving all checkpoints intact for the next run. |
| Steady-state pacing | REQUEST_DELAY between calls β 0.2s with a token, 1.5s without. |
| Symptom | Likely cause | Resolution |
|---|---|---|
| 401 Unauthorized / 403 Forbidden in the log | Token missing, mistyped, expired or revoked | Confirm the HF_TOKEN environment variable is visible to the process (echo $HF_TOKEN / echo %HF_TOKEN%). Issue a fresh read token and retry. |
| Frequent HTTP 429 lines, run appears stalled | Hitting the account's 5-minute quota | This is normal and self-correcting β the script waits and continues. To reduce frequency, raise REQUEST_DELAY to 0.5, or move to a higher-tier organisation token. |
get_dataset_leaderboard is missing / AttributeError |
Outdated huggingface_hub |
pip install -U huggingface_hub and re-run the Appendix B check. |
| "No orgs discovered yet β run steps 2/3 first" | Stage 5 executed with no model/dataset checkpoints present (typical after a FRESH_START that was interrupted early) | Let stages 2 and 3 complete, then re-run. Expected behaviour, not an error. |
| "No benchmarks found (or the call failed after retries)" | Transient API failure, or the token cannot read the benchmark listings | Re-run. If it persists, verify network access to huggingface.co and that the token is valid. |
Some benchmarks listed in 1_leaderboards_skipped.txt |
Those datasets have no leaderboard, or it could not be parsed | Expected and harmless β benchmarks without a leaderboard simply have no data. Review the file as a data-quality note; delete a line to force a retry. |
| Excel truncation warning printed | A source exceeded the 1,000,000-row Excel cap | None needed. Use the CSV copy, which is complete. For a full dataset of that size, load the CSV or JSONL into a database instead of Excel. |
openpyxl / IllegalCharacterError |
Control characters in Hub-provided text | Already handled β the exporter sanitises illegal characters before writing. If it recurs, confirm openpyxl>=3.1.2 is installed. |
AttributeError: 'DataFrame' object has no attribute 'map' |
pandas older than 2.1 | pip install -U "pandas>=2.1.0". |
| Workbook is very slow to open / huge file size | Hundreds of columns and many rows | Expected for this dataset shape. Open the CSVs in a notebook or BI tool instead of Excel for day-to-day analysis. |
| Machine slows down or the process is killed on a long run | Accumulating in-memory "already done" key sets, or OS memory pressure | Stop and re-run β it will resume. Avoid running other heavy workloads concurrently; see Section 15 for the memory note. |
| Run finished but no new rows appeared | Window is genuinely empty, or the stage was already at full coverage | Check the banner's scan windows. The range-only workbook will contain a "No New Data" placeholder sheet explaining that no rows were fetched in the period. |
| Unexpectedly huge backfill after a code change | FRESH_START left at True, or pipeline_state.json deleted |
Verify FRESH_START = False. If the state file was lost, the next clean run re-establishes it. |
| Cadence | Rationale |
|---|---|
| Weekly | The natural rhythm for this dataset. Each run reports exactly what appeared since the last successful run, and the range-only workbook filename states that period explicitly β ideal for a Monday-morning briefing pack. |
| Monthly | Acceptable if the client only needs periodic snapshots. Runs take longer because the window is wider, and the chance of hitting rate limits is higher. |
| More often than daily | Not advisable. The Hub's creation volume in a sub-day window is small, and repeated runs consume quota without adding much data. |
pipeline_state.json exists, runs are incremental and much shorter β suitable for unattended scheduling.# Run every Monday at 03:00, logging to a dated file
0 3 * * 1 cd /opt/hf-intelligence && \
HF_TOKEN="hf_..." /opt/hf-intelligence/.venv/bin/python \
hf_intelligence_pipeline.py >> logs/run_$(date +\%F).log 2>&1
Prefer putting the token in a root-owned environment file that cron sources, rather than inline as shown, so it does not appear in the crontab or process list.
C:\hf-intelligence\.venv\Scripts\python.exe; arguments hf_intelligence_pipeline.py; start-in C:\hf-intelligence.Finished β full run completed successfully. means the window advanced. Finished β run was interruptedβ¦ means it did not, and the next run will cover the same period again.5_orgs_completed.txt and 1_leaderboards_skipped.txt as audit evidence of coverage.Back up the whole output folder on the same cadence as the pipeline β the master checkpoints are the accumulated research asset and are not reproducible without re-crawling. Because the crawl is incremental, losing pipeline_state.json and the coverage markers alongside the data means a full 2024-onwards re-crawl.
HF_TOKEN = "hf_β¦"). Anyone holding a copy of this file holds that credential. Remediation:
huggingface.co/settings/tokens and revoke the token currently in the file.HF_TOKEN environment variable (Section 5, Step 5).| Rule | Reason |
|---|---|
| Use a read-scoped token | The pipeline only reads public data. A read token cannot modify or delete anything even if it leaks. |
| Never commit the token to Git | It stays recoverable in history forever. If the project is version-controlled, add the venv, output folder and any .env file to .gitignore. |
| One token per deployment | Lets you revoke a single environment without disturbing others, and makes usage traceable in the account's activity log. |
| Rotate on a schedule (e.g. every 90 days) and on staff change | Limits the blast radius of any undisclosed exposure. |
| Prefer an organisation token for production | Higher rate limits, and the token belongs to the organisation rather than an individual β so it survives staff turnover. |
| Restrict file-system access to the output folder | It contains a complete profile of third-party organisations. Apply the client's normal data-handling policy. |
REQUEST_DELAY to zero.The following behaviours are by design or inherent to the approach. They are documented so the client can plan around them rather than treat them as defects.
| Limitation | Detail and recommended mitigation |
|---|---|
| Leaderboards are snapshotted once per benchmark | Once a benchmark has been fetched, it is permanently marked done, so scores added to that benchmark later are not picked up. Mitigation: to refresh a benchmark, remove its dataset ID from 1_leaderboards.jsonl (or delete the whole file to refresh everything) and re-run. Document the intended refresh policy with the client up front. |
| Organisation footprint is captured once per organisation | An organisation already recorded in 5_orgs_completed.txt is not re-profiled, so repositories it publishes later will not appear in source 5 β though genuinely new repositories do appear in sources 2 and 3. Mitigation: for a periodic re-profile, delete the chosen org names from 5_orgs_completed.txt (and optionally their rows from the two 5_org_*.jsonl files) before a run. |
| Stage 5 has no upper bound on scope | Organisation discovery is dynamic, so a wide date window can surface tens of thousands of namespaces, each requiring one or two API calls. The stage is designed to be interruptible and resumable; split the work across several sittings if necessary. |
| In-memory "already done" key sets grow during a run | Each stage holds the set of processed IDs in memory while it runs. On very large crawls this consumes a noticeable amount of RAM. Mitigation: 16 GB RAM recommended; stopping and restarting rebuilds the sets from the checkpoint files, which releases memory. |
| Model-info cache is per-run only | Leaderboard enrichment caches model_info responses in memory and does not persist them between runs, so a re-run of stage 1 repeats those lookups. Acceptable given the stage is normally run once per benchmark. |
| Wide, sparse tables in the export | Because records carry different fields, the flat export has many columns and many blanks. Mitigation: prefer the CSVs over Excel, load into a database for repeated querying, and define analysis views rather than working with the raw wide table. |
| Excel is not a delivery format for full volumes | Sheets are capped at 1,000,000 rows by the script (Excel's own limit is 1,048,576). CSV and JSONL always hold the complete data. Agree with the client in advance which format is authoritative for downstream use. |
| No automated alerting | Failures are reported in the console and log file only. Mitigation: add a wrapper script that checks the final log line and sends an email/Teams message on an unexpected exit code. |
| No schema versioning on the checkpoints | Field sets can shift if the Hub changes its API. Existing checkpoints are never rewritten, so a change affects only new rows. Note the date of any Hub API change in the run log. |
| Requirement | How to implement |
|---|---|
| Move output to another drive | Change OUTPUT_DIR to an absolute path, e.g. r"D:\hf_intelligence_data". |
| Change the historical backfill start | Change _ORIGINAL_ANCHOR_START_DATE, then run once with FRESH_START = True and set it back to False. |
| Re-collect a specific week only | Temporarily comment out stages 1 and 5 in the __main__ block and set both full-coverage states to the required start date, or delete them β see Section 9. |
| Speed up on a high-limit plan | Lower REQUEST_DELAY to 0.1 and re-run; monitor for 429s. |
| Add a new data source | Copy the shape of an existing stage: a dedicated fetch_β¦ function, its own JSONL checkpoint, a done-key field, and an entry in the export's sources list. |
| Push results to a database or dashboard | Load the exported CSVs (or the JSONL directly) with a downstream job. The checkpoint files are append-only, so a simple incremental load on id works well. |
Ctrl+C.pip install -U huggingface_hub requests pandas openpyxl.1_leaderboards.jsonl and 2_new_models.jsonl tails for unexpected empty fields.REQUEST_DELAY).python --version) and output of pip freeze.pipeline_state.json, the cursor files and the coverage state files.1_leaderboards_skipped.txt.| Stage | Endpoint / method | Purpose |
|---|---|---|
| 1 | HfApi.list_datasets(benchmark=True) | Enumerate every benchmark dataset on the Hub. |
| 1 | GET /api/datasets/{dataset_id}/leaderboardvia HfApi.get_dataset_leaderboard() | Fetch the ranked model scores for one benchmark. |
| 1 | HfApi.model_info(model_id, securityStatus=True) | Full metadata dump for each ranked model, including its card data. |
| 2 | GET /api/models?sort=createdAt&direction=-1&limit=100&full=true&cardData=true | Paginated listing of all models, newest first, with full metadata. |
| 3 | GET /api/datasets?sort=createdAt&direction=-1&limit=100&full=true&cardData=true | Paginated listing of all datasets, newest first, with full metadata. |
| 4 | GET /api/daily_papers?date=YYYY-MM-DD | The curated paper feed for one calendar day. |
| 5 | GET /api/models?author={org} and GET /api/datasets?author={org} | Complete historical catalogue for one organisation. |
Stages 2, 3 and 5 follow the Hub's Link response header (rel="next") to walk pages. The next-page URL is saved after every page, which is what makes mid-crawl resumption possible.
HfApi Python client β huggingface.co/docs/huggingface_hub/package_reference/hf_apiSave this as preflight_check.py in the project folder and run it with the virtual environment active before the first real run and after any library upgrade. It validates the token, network access, library capabilities and date resolution without collecting any data.
"""
preflight_check.py -- validates the environment for hf_intelligence_pipeline.py
Run: python preflight_check.py
"""
import os, sys, json
from datetime import datetime, timezone
print("=" * 68)
print("HuggingFace Intelligence Pipeline -- pre-flight check")
print("=" * 68)
ok = True
# 1. Python version ---------------------------------------------------------
major, minor = sys.version_info[:2]
print(f"\n[1] Python version: {major}.{minor}")
if (major, minor) < (3, 9):
print(" FAIL - Python 3.9 or newer is required.")
ok = False
else:
print(" OK")
# 2. Packages ---------------------------------------------------------------
print("\n[2] Required packages")
try:
import pandas as pd
print(f" pandas {pd.__version__}")
if pd.__version__ < "2.1":
print(" FAIL - pandas 2.1+ is required (DataFrame.map).")
ok = False
except ImportError:
print(" FAIL - pandas not installed. Run: pip install -r requirements.txt")
ok = False
try:
import requests
print(f" requests {requests.__version__}")
except ImportError:
print(" FAIL - requests not installed.")
ok = False
try:
import openpyxl
print(f" openpyxl {openpyxl.__version__}")
except ImportError:
print(" FAIL - openpyxl not installed.")
ok = False
try:
import huggingface_hub
from huggingface_hub import HfApi
print(f" huggingface_hub {huggingface_hub.__version__}")
except ImportError:
print(" FAIL - huggingface_hub not installed.")
ok = False
# 3. Library capability -----------------------------------------------------
print("\n[3] Library capabilities")
api = HfApi()
for name in ("get_dataset_leaderboard", "list_datasets", "model_info"):
if hasattr(api, name):
print(f" OK HfApi.{name}")
else:
print(f" FAIL HfApi.{name} is missing -- upgrade huggingface_hub.")
ok = False
# 4. Token ------------------------------------------------------------------
print("\n[4] Authentication")
token = os.environ.get("HF_TOKEN", "").strip()
if not token:
print(" FAIL - HF_TOKEN environment variable is not set (or is empty).")
print(" See documentation Section 14 for how to set it safely.")
ok = False
else:
print(f" Token found (length {len(token)}, prefix {token[:6]}...)")
try:
who = HfApi(token=token).whoami()
print(f" OK Authenticated as: {who.get('name')} ({who.get('type')})")
except Exception as e:
print(f" FAIL - token rejected by the Hub: {e}")
ok = False
# 5. Network reachability ---------------------------------------------------
print("\n[5] Network reachability")
try:
r = requests.get("https://huggingface.co/api/models?limit=1",
headers=({"Authorization": f"Bearer {token}"} if token else {}),
timeout=15)
print(f" OK GET /api/models returned HTTP {r.status_code}")
except Exception as e:
print(f" FAIL - cannot reach huggingface.co: {e}")
ok = False
# 6. Output directory -------------------------------------------------------
print("\n[6] Output directory")
out_dir = "HuggingFace_Intelligence_Output"
try:
os.makedirs(out_dir, exist_ok=True)
probe = os.path.join(out_dir, ".write_test")
with open(probe, "w") as f:
f.write("ok")
os.remove(probe)
print(f" OK '{out_dir}' exists and is writable")
except Exception as e:
print(f" FAIL - cannot write to '{out_dir}': {e}")
ok = False
# 7. Date window ------------------------------------------------------------
print("\n[7] Date window resolution")
state_path = os.path.join(out_dir, "pipeline_state.json")
if os.path.exists(state_path):
try:
with open(state_path, encoding="utf-8") as f:
state = json.load(f)
print(f" Previous successful run through: {state.get('last_successful_run_end_date')}")
except Exception as e:
print(f" WARN - could not read {state_path}: {e}")
else:
print(" No pipeline_state.json yet -> first run will backfill from 2024-06-01.")
print(f" END_DATE for the next run will be: {datetime.now(timezone.utc).isoformat()}")
# 8. Disk space -------------------------------------------------------------
print("\n[8] Free disk space")
try:
import shutil
free_gb = shutil.disk_usage(os.path.abspath(out_dir)).free / (1024 ** 3)
print(f" {free_gb:.1f} GB free on the output drive")
if free_gb < 20:
print(" WARN - less than 20 GB free; a full crawl needs substantial space.")
except Exception as e:
print(f" WARN - could not determine disk space: {e}")
print("\n" + "=" * 68)
print("RESULT:", "ALL CHECKS PASSED - safe to run the pipeline." if ok
else "ONE OR MORE CHECKS FAILED - fix the items marked FAIL above.")
print("=" * 68)
sys.exit(0 if ok else 1)
| Term | Definition |
|---|---|
| Anchor date | The historical backstop (2024-06-01) used for the very first crawl or after a full wipe. |
| Benchmark | A dataset repo on the Hub that carries a leaderboard ranking models by evaluation score. |
base_model | The upstream model a model was derived from, as declared in its model card metadata. |
| Checkpoint | A JSONL file that accumulates collected records so a run can be stopped and resumed safely. |
| Coverage state | A file recording the date through which a full crawl has already been confirmed, so future runs need not repeat it. |
| Delta | The rows added during a single run, as distinct from the cumulative master history. |
| Done-keys | The set of identifiers already present in a checkpoint, used to skip work on re-runs. |
| Hub | The Hugging Face model, dataset and Space hosting platform. |
| JSONL | JSON Lines β one complete JSON object per line. Tolerant of records with differing field sets. |
| Namespace / org | The owner portion of a repository ID. Used as the company-level identifier throughout the outputs. |
| Rate limit | The Hub's cap on API requests, counted over fixed five-minute windows. |
| Resume cursor | The saved next-page URL that lets a crawl restart mid-way instead of from the top. |
| Run window | The START_DATE β END_DATE period a given execution covers. |
Use this list to confirm the pipeline is genuinely ready for client operation. Each item should be verified by the client, not assumed.
| # | Item | Verified by / date |
|---|---|---|
| 1 | Token rotated. The token present in the delivered script has been revoked and replaced; the replacement is supplied via environment variable only. | |
| 2 | Python 3.9+ installed and a virtual environment created on the target machine. | |
| 3 | pip install -r requirements.txt completed without error. | |
| 4 | preflight_check.py (Appendix B) reports ALL CHECKS PASSED. | |
| 5 | At least 20 GB free disk space confirmed on the output drive. | |
| 6 | OUTPUT_DIR set to the agreed production location. | |
| 7 | FRESH_START confirmed as False for normal operation. | |
| 8 | First full run executed to completion; pipeline_state.json created. | |
| 9 | Master workbook and range-only workbook both opened successfully and sheet contents spot-checked against the Hub. | |
| 10 | Second (incremental) run executed; banner confirmed it scanned only the new period. | |
| 11 | Scheduling configured (cron / Task Scheduler) with output redirected to a log file. | |
| 12 | Backup routine for the output folder agreed and in place. | |
| 13 | Refresh policy agreed for leaderboards and organisation footprints (Section 15). | |
| 14 | Support/escalation contact and the information-to-capture list (Section 16.3) shared with the client's operations team. | |
| 15 | Client has received this document and confirmed the run instructions were followed unaided at least once. |