How to Use Hugging Face Skills for Model Training and Evaluation: A Complete Guide

The VoltAgent awesome-agent-skills repository provides self-contained Agent Skills that enable LLM agents to execute full ML workflows—including SFT/DPO/GRPO training with TRL, GGUF export, and vLLM-powered evaluation—through standardized JSON invocations.

The awesome-agent-skills repository curates official Hugging Face Agent Skills that expose the complete machine learning lifecycle to LLM-powered agents. These skills allow you to orchestrate data handling, model training, evaluation, and experiment tracking without leaving your agentic chat interface.

Overview of Hugging Face Agent Skills for ML Workflows

The repository indexes skills covering every phase of the ML pipeline. In README.md (lines 322-336), you will find the complete catalog of Hugging Face skills mapped to specific workflow phases.

Data Management Skills

  • hugging-face-dataset-viewer – Browse, filter, and query Hugging Face datasets via the Dataset Viewer API using commands like hf dataset view …
  • hugging-face-datasets – Create, configure, and query datasets using SQL-style operations via hf dataset create …

Model Training Skills

  • hugging-face-model-trainer – Executes SFT, DPO, and GRPO training using the TRL library, with automatic export to GGUF format. The typical entry point is hf train …
  • hugging-face-vision-trainer – Specialized trainer for image models with the command hf vision-train …

Evaluation and Orchestration Skills

  • hugging-face-evaluation – Runs model benchmarks using vLLM, lighteval, and custom evaluation tables via hf eval …
  • hugging-face-trackio – Provides real-time dashboards for metrics, hyper-parameters, and artifacts accessible through hf track …
  • hugging-face-jobs – Submits arbitrary Python scripts or container jobs to Hugging Face compute resources using hf job submit …

Architecture and Invocation Pattern

All skills follow a standardized invocation pattern. When an LLM agent decides to perform an action, it constructs a JSON payload specifying the skill identifier and inputs.

The architectural flow implemented in the VoltAgent platform follows four stages:

  1. Agent → Skill request – The LLM selects the appropriate Hugging Face skill based on user intent
  2. Skill runtime – The skill container authenticates with the HF Hub using a secret token stored in the agent's runtime environment
  3. Execution – The skill invokes the HF REST API or Python SDK to perform the requested action (e.g., hf train)
  4. Feedback – Progress streams back to the agent, returning a structured response containing model IDs, evaluation scores, and track.io dashboard URLs

The standard JSON invocation structure is:

{
  "skill": "huggingface/hugging-face-model-trainer",
  "inputs": {
    "dataset": "my-org/my-dataset",
    "model": "meta-llama/Meta-Llama-3-8B",
    "training_type": "sft",
    "epochs": 3,
    "learning_rate": 5e-5
  }
}

Practical Code Examples for Model Training

Fine-Tuning with the Model Trainer

To invoke the hugging-face-model-trainer skill from a Claude-based agent using JavaScript:

import { Agent } from '@voltagent/core'

// Create an agent that has access to Hugging Face skills
const agent = new Agent({
  apiKey: process.env.VOLTAGENT_API_KEY,
  enabledSkills: ['huggingface/hugging-face-model-trainer']
})

// Prompt the user to fine-tune a model
const response = await agent.run(`
  Fine-tune Llama-3-8B on my dataset "my-org/my-dataset" for 2 epochs.
`)

console.log('Trainer output:', response.result)
// => {
//   modelId: "hf:meta-llama/Meta-Llama-3-8B-finetuned-2024-04-22",
//   evalScore: 0.87,
//   dashboardUrl: "https://track.io/dashboard/…"
// }

This executes TRL-based training with support for SFT, DPO, and GRPO methods, automatically exporting to GGUF format upon completion.

Running Evaluation Benchmarks

Use the hugging-face-evaluation skill to benchmark fine-tuned models via vLLM and lighteval:

from voltagent import Agent

agent = Agent(
    api_key=os.getenv("VOLTAGENT_API_KEY"),
    enabled_skills=["huggingface/hugging-face-evaluation"]
)

prompt = """
Evaluate the fine-tuned model "hf:my-org/llama3-finetuned" on the "lambada" benchmark.
"""

result = agent.run(prompt)

print("Eval table URL:", result["tableUrl"])
print("Average score:", result["metrics"]["accuracy"])

Submitting Custom Training Jobs

For arbitrary training scripts, use the hugging-face-jobs skill directly or via CLI:

hf job submit \
  --script train.py \
  --requirements requirements.txt \
  --env HF_TOKEN=$HF_TOKEN \
  --gpu a100 \
  --timeout 12h

Wrapped as an Agent Skill call, this becomes:

{
  "skill": "huggingface/hugging-face-jobs",
  "inputs": {
    "script_path": "train.py",
    "requirements": "requirements.txt",
    "gpu_type": "a100",
    "timeout": "12h"
  }
}

Configuration and Discovery Files

The repository structure enables skill discovery through metadata files:

  • README.md (lines 322-336) – Contains the comprehensive catalog of Hugging Face skills and links to their definitions on officialskills.sh
  • opencode.json – Metadata used by the VoltAgent platform to discover and load skill bundles

These files provide the index that agents use to resolve skill identifiers like huggingface/hugging-face-model-trainer to their executable implementations.

Summary

  • Hugging Face skills in the awesome-agent-skills repository cover the complete ML lifecycle from data management to deployment
  • The hugging-face-model-trainer skill supports SFT, DPO, and GRPO training via TRL with automatic GGUF export
  • Evaluation leverages vLLM and lighteval through the hugging-face-evaluation skill for comprehensive benchmarking
  • Skills are invoked via JSON payloads containing skill identifiers and parameter inputs, executed through a four-phase agent runtime
  • Real-time experiment tracking is available through track.io integration via the hugging-face-trackio skill

Frequently Asked Questions

What training methods does the Hugging Face Model Trainer skill support?

According to the source code analysis, the hugging-face-model-trainer skill supports SFT (Supervised Fine-Tuning), DPO (Direct Preference Optimization), and GRPO (Generalized Reward-Penalty Optimization) training methods. It utilizes the TRL library under the hood and automatically exports trained models to GGUF format for efficient inference.

How do Hugging Face skills authenticate with the Hugging Face Hub?

The skills authenticate using a secret token stored in the agent's runtime environment. When the skill container executes, it retrieves this token to authorize REST API calls and SDK operations against the HF Hub, enabling secure access to private datasets and model repositories.

Can I use these skills for computer vision model training?

Yes. The repository includes hugging-face-vision-trainer, a specialized skill for image model training. While the standard hugging-face-model-trainer handles text-based LLMs, the vision trainer provides optimized pipelines for image classification and other computer vision tasks.

Where can I find the complete list of available Hugging Face skills?

The complete catalog is documented in README.md starting at line 322, which maps each skill to its phase in the ML workflow (data, training, evaluation, etc.). The file also links to officialskills.sh for detailed skill definitions, while opencode.json provides machine-readable metadata for programmatic discovery.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →