Video summary
How I automate my own job at Hugging Face using agents — Niels Rogge, Hugging Face
Main summary
Key takeaways
Technological concepts & product/platform features highlighted
Problem: poor discoverability of research artifacts
- Many researchers release model weights/data artifacts outside Hugging Face (e.g., Google Drive, GitHub releases, Dropbox, Zenodo).
- This makes it harder for others to discover, reproduce, and reuse work.
- Community outreach goal: encourage researchers to publish checkpoints, models, and datasets directly on Hugging Face.
Why Hugging Face helps discoverability
- Paper pages (linked from arXiv) show linked artifacts on the right side (e.g., models/datasets).
- Metadata tags/filters on models and datasets improve search (e.g., task type, language, compatible library).
- Model cards / dataset cards improve documentation and usability, and Hugging Face supports tooling to upload/download artifacts.
Review / guide / tutorial-like recommendations (explicit)
- Recommends Anthropic’s blog post: “Building effective agents”
- Suggests starting simple and preferring deterministic workflows over agentic loops when possible.
- Recommends Hamel Husain’s “LLM Evils FAQ”
- Focuses on evaluating agents to avoid producing low-quality “slop” outputs.
Automated system: “agentic” outreach workflow to scale community science
Core workflow being automated
- Identify a paper and find its GitHub URL
- Read the GitHub README
- Check whether the work is already present on Hugging Face
- If missing/insufficient:
- open a GitHub issue requesting Hugging Face release of weights/datasets
- otherwise open a PR to improve model cards/data set cards and add missing metadata tags
- Follow up with the author
Initial approach: deterministic workflow (not fully autonomous)
- Implemented as a step-by-step pipeline using LLM APIs with no agent framework, emphasizing predictability and control.
- Visualized as a diagram/pipeline using Excalidraw MCP server in Cursor.
Scheduling & orchestration
- Runs nightly via a cron job implemented as a Python script calling LLM APIs.
- Uses GitHub Actions to manage cron jobs.
- Purpose: process hundreds of arXiv papers daily.
Observability / tracing
- Uses LangFuse for tracing and observability:
- inputs/outputs, prompts, cost, latency, and what the LLM is doing.
Follow-up automation: moving to fully autonomous agents
Motivation
- Creating many GitHub issues generates lots of unread notifications due to replies—creating manual workload.
Second system: fully autonomous follow-up
- Uses Claude agents SDK (initially).
- Motivation includes an Anthropic workshop suggesting agents may now outperform workflows due to improved model capability.
- Also notes that Cursor can replace thousands of lines of code with a small agent “skill,” implying reduced custom logic.
Architecture & deployment details for the follow-up agent
Agent SDK + model serving
- Built with Claude agents SDK as a Python SDK.
- Later updated to use GLM 5.2 via Hugging Face Inference Providers (an OpenAI/Anthropic-compatible wrapper over multiple providers such as Together AI, Fireworks, Cerebras, etc.).
Tools
- Uses Modal deployment.
- Employs a Bash/terminal tool plus a Hugging Face CLI skill to run Hugging Face CLI commands.
- Can comment on GitHub follow-ups.
- Posts final results to Slack (Slack updates are part of the agent’s output loop).
Parallelism / batch processing
- Modal’s batch processing runs many containerized agent loops in parallel:
- “one container = one agent loop processing one GitHub issue”
- Suitable for overnight/background workloads.
Invocation structure
- Uses a skill (called
processunder Modal in Cursor context) that:- invokes an agent loop (e.g., Composer 2.5 as described)
- which then invokes other agents
- then posts results to Slack
Outcomes / qualitative results (implied “evaluation”)
Scale and interaction
- Generates hundreds of GitHub issues nightly.
- Reported as “mostly win-win”: only two negative comments mentioned; most responses accept publishing models/weights on Hugging Face.
Examples of impact
- Companies/research groups migrating artifacts to Hugging Face due to agent outreach:
- Paddle OCR moved OCR models
- outreach examples to Apple and Google DeepMind
- “Fun” automations:
- auto-completing model card templates (referencing “model cards for model reporting”)
- generating model card text that humorously includes the speaker
- High-engagement case:
- “Tiny Recursive Models” issue was upvoted by 60+ people, contributing to release on Hugging Face.
Anti-spam / quality control stance
- To avoid “slop,” emphasizes agent evaluation and references the LLM Evils FAQ.
- Also chooses not to explicitly disclose the assistant is a bot, arguing people might close the issue otherwise.
Additional efforts beyond GitHub outreach
“Daily Papers” Twitter/X account
- Uses the same workflow “behind the scenes.”
- Posts research papers/artifacts every ~4 hours or upon new notable Hugging Face releases.
- Mentions a Gemini component for selecting visuals for tweets.
Revival of “Papers With Code”
- Reintroduces a benchmarking/educational platform at paperswithcode.co.
- Includes benchmark listings (e.g., OCR benches) and educational content (e.g., concepts like mixed training, policy distillation).
Main speakers / sources
- Main speaker: Niels Rogge (Hugging Face; community science team; “Niels from Belgium”)
Referenced sources/tools
- Anthropic blog: “Building effective agents”
- Hamel Husain: “LLM Evils FAQ”
- LangFuse (tracing/observability)
- Claude agents SDK (initial follow-up agent framework)
- Modal (deployment + batch parallelism)
- Hugging Face Inference Providers (GLM 5.2 access)
- GitHub Actions (cron job management)
- Composer 2.5, Cursor (mentioned in context of agent skills/architecture)
- Gemini (visual selection for Daily Papers)
- Hugging Face CLI skill (tooling for agent actions)