{"id":5944,"plugin_id":"plugin_asdk_app_6939e86417648191b7bda087d872685b","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T22:47:17.020Z","digest":"aa464ed9ecd86db99e9be64fc8e716ddc105c9e7732608d7396e0c3117dcc2d8","against":null,"payload":{"name":"huggingface-vision-trainer","description":"Trains and fine-tunes vision models for object detection (D-FINE, RT-DETR v2, DETR, YOLOS), image classification (timm models — MobileNetV3, MobileViT, ResNet, ViT/DINOv3 — plus any Transformers classifier), and SAM/SAM2 segmentation using Hugging Face Transformers on Hugging Face Jobs cloud GPUs. Covers COCO-format dataset preparation, Albumentations augmentation, mAP/mAR evaluation, accuracy metrics, SAM segmentation with bbox/point prompts, DiceCE loss, hardware selection, cost estimation, Trackio monitoring, and Hub persistence. Use when users mention training object detection, image classification, SAM, SAM2, segmentation, image matting, DETR, D-FINE, RT-DETR, ViT, timm, MobileNet, ResNet, bounding box models, or fine-tuning vision models on Hugging Face Jobs.","included_files":[{"relative_path":"references/finetune_sam2_trainer.md","size_in_bytes":6432},{"relative_path":"references/hub_saving.md","size_in_bytes":16949},{"relative_path":"references/image_classification_training_notebook.md","size_in_bytes":11075},{"relative_path":"references/object_detection_training_notebook.md","size_in_bytes":29917},{"relative_path":"references/reliability_principles.md","size_in_bytes":9487},{"relative_path":"references/timm_trainer.md","size_in_bytes":3547},{"relative_path":"scripts/dataset_inspector.py","size_in_bytes":31851},{"relative_path":"scripts/estimate_cost.py","size_in_bytes":7400},{"relative_path":"scripts/image_classification_training.py","size_in_bytes":13385},{"relative_path":"scripts/object_detection_training.py","size_in_bytes":27371},{"relative_path":"scripts/sam_segmentation_training.py","size_in_bytes":13946}],"skill_md_contents":"---\nname: huggingface-vision-trainer\ndescription: Trains and fine-tunes vision models for object detection (D-FINE, RT-DETR v2, DETR, YOLOS), image classification (timm models — MobileNetV3, MobileViT, ResNet, ViT/DINOv3 — plus any Transformers classifier), and SAM/SAM2 segmentation using Hugging Face Transformers on Hugging Face Jobs cloud GPUs. Covers COCO-format dataset preparation, Albumentations augmentation, mAP/mAR evaluation, accuracy metrics, SAM segmentation with bbox/point prompts, DiceCE loss, hardware selection, cost estimation, Trackio monitoring, and Hub persistence. Use when users mention training object detection, image classification, SAM, SAM2, segmentation, image matting, DETR, D-FINE, RT-DETR, ViT, timm, MobileNet, ResNet, bounding box models, or fine-tuning vision models on Hugging Face Jobs.\n---\n\n# Vision Model Training on Hugging Face Jobs\n\nTrain object detection, image classification, and SAM/SAM2 segmentation models on managed cloud GPUs. No local GPU setup required—results are automatically saved to the Hugging Face Hub.\n\n## When to Use This Skill\n\nUse this skill when users want to:\n- Fine-tune object detection models (D-FINE, RT-DETR v2, DETR, YOLOS) on cloud GPUs or local\n- Fine-tune image classification models (timm: MobileNetV3, MobileViT, ResNet, ViT/DINOv3, or any Transformers classifier) on cloud GPUs or local\n- Fine-tune SAM or SAM2 models for segmentation / image matting using bbox or point prompts\n- Train bounding-box detectors on custom datasets\n- Train image classifiers on custom datasets\n- Train segmentation models on custom mask datasets with prompts\n- Run vision training jobs on Hugging Face Jobs infrastructure\n- Ensure trained vision models are permanently saved to the Hub\n\n## Related Skills\n\n- **`hugging-face-jobs`** — General HF Jobs infrastructure: token authentication, hardware flavors, timeout management, cost estimation, secrets, environment variables, scheduled jobs, and result persistence. **Refer to the Jobs skill for any non-training-specific Jobs questions** (e.g., \"how do secrets work?\", \"what hardware is available?\", \"how do I pass tokens?\").\n- **`hugging-face-model-trainer`** — TRL-based language model training (SFT, DPO, GRPO). Use that skill for text/language model fine-tuning.\n\n## Local Script Execution\n\nHelper scripts use PEP 723 inline dependencies. Run them with `uv run`:\n```bash\nuv run scripts/dataset_inspector.py --dataset username/dataset-name --split train\nuv run scripts/estimate_cost.py --help\n```\n\n## Prerequisites Checklist\n\nBefore starting any training job, verify:\n\n### Account & Authentication\n- Hugging Face Account with [Pro](https://hf.co/pro), [Team](https://hf.co/enterprise), or [Enterprise](https://hf.co/enterprise) plan (Jobs require paid plan)\n- Authenticated login: Check with `hf_whoami()` (tool) or `hf auth whoami` (terminal)\n- Token has **write** permissions\n- **MUST pass token in job secrets** — see directive #3 below for syntax (MCP tool vs Python API)\n\n### Dataset Requirements — Object Detection\n- Dataset must exist on Hub\n- Annotations must use the `objects` column with `bbox`, `category` (and optionally `area`) sub-fields\n- Bboxes can be in **xywh (COCO)** or **xyxy (Pascal VOC)** format — auto-detected and converted\n- Categories can be **integers or strings** — strings are auto-remapped to integer IDs\n- `image_id` column is **optional** — generated automatically if missing\n- **ALWAYS validate unknown datasets** before GPU training (see Dataset Validation section)\n\n### Dataset Requirements — Image Classification\n- Dataset must exist on Hub\n- Must have an **`image` column** (PIL images) and a **`label` column** (integer class IDs or strings)\n- The label column can be `ClassLabel` type (with names) or plain integers/strings — strings are auto-remapped\n- Common column names auto-detected: `label`, `labels`, `class`, `fine_label`\n- **ALWAYS validate unknown datasets** before GPU training (see Dataset Validation section)\n\n### Dataset Requirements — SAM/SAM2 Segmentation\n- Dataset must exist on Hub\n- Must have an **`image` column** (PIL images) and a **`mask` column** (binary ground-truth segmentation mask)\n- Must have a **prompt** — either:\n  - A **`prompt` column** with JSON containing `{\"bbox\": [x0,y0,x1,y1]}` or `{\"point\": [x,y]}`\n  - OR a dedicated **`bbox`** column with `[x0,y0,x1,y1]` values\n  - OR a dedicated **`point`** column with `[x,y]` or `[[x,y],...]` values\n- Bboxes should be in **xyxy** format (absolute pixel coordinates)\n- Example dataset: `merve/MicroMat-mini` (image matting with bbox prompts)\n- **ALWAYS validate unknown datasets** before GPU training (see Dataset Validation section)\n\n### Critical Settings\n- **Timeout must exceed expected training time** — Default 30min is TOO SHORT. See directive #6 for recommended values.\n- **Hub push must be enabled** — `push_to_hub=True`, `hub_model_id=\"username/model-name\"`, token in `secrets`\n\n## Dataset Validation\n\n**Validate dataset format BEFORE launching GPU training to prevent the #1 cause of training failures: format mismatches.**\n\n**ALWAYS validate for** unknown/custom datasets or any dataset you haven't trained with before. **Skip for** `cppe-5` (the default in the training script).\n\n### Running the Inspector\n\n**Option 1: Via HF Jobs (recommended — avoids local SSL/dependency issues):**\n```python\nhf_jobs(\"uv\", {\n    \"script\": \"path/to/dataset_inspector.py\",\n    \"script_args\": [\"--dataset\", \"username/dataset-name\", \"--split\", \"train\"]\n})\n```\n\n**Option 2: Locally:**\n```bash\nuv run scripts/dataset_inspector.py --dataset username/dataset-name --split train\n```\n\n**Option 3: Via `HfApi().run_uv_job()` (if hf_jobs MCP unavailable):**\n```python\nfrom huggingface_hub import HfApi\napi = HfApi()\napi.run_uv_job(\n    script=\"scripts/dataset_inspector.py\",\n    script_args=[\"--dataset\", \"username/dataset-name\", \"--split\", \"train\"],\n    flavor=\"cpu-basic\",\n    timeout=300,\n)\n```\n\n### Reading Results\n\n- **`✓ READY`** — Dataset is compatible, use directly\n- **`✗ NEEDS FORMATTING`** — Needs preprocessing (mapping code provided in output)\n\n## Automatic Bbox Preprocessing\n\nThe object detection training script (`scripts/object_detection_training.py`) automatically handles bbox format detection (xyxy→xywh conversion), bbox sanitization, `image_id` generation, string category→integer remapping, and dataset truncation. **No manual preprocessing needed** — just ensure the dataset has `objects.bbox` and `objects.category` columns.\n\n## Training workflow\n\nCopy this checklist and track progress:\n\n```\nTraining Progress:\n- [ ] Step 1: Verify prerequisites (account, token, dataset)\n- [ ] Step 2: Validate dataset format (run dataset_inspector.py)\n- [ ] Step 3: Ask user about dataset size and validation split\n- [ ] Step 4: Prepare training script (OD: scripts/object_detection_training.py, IC: scripts/image_classification_training.py, SAM: scripts/sam_segmentation_training.py)\n- [ ] Step 5: Save script locally, submit job, and report details\n```\n\n**Step 1: Verify prerequisites**\n\nFollow the Prerequisites Checklist above.\n\n**Step 2: Validate dataset**\n\nRun the dataset inspector BEFORE spending GPU time. See \"Dataset Validation\" section above.\n\n**Step 3: Ask user preferences**\n\nALWAYS use the AskUserQuestion tool with option-style format:\n\n```python\nAskUserQuestion({\n    \"questions\": [\n        {\n            \"question\": \"Do you want to run a quick test with a subset of the data first?\",\n            \"header\": \"Dataset Size\",\n            \"options\": [\n                {\"label\": \"Quick test run (10% of data)\", \"description\": \"Faster, cheaper (~30-60 min, ~$2-5) to validate setup\"},\n                {\"label\": \"Full dataset (Recommended)\", \"description\": \"Complete training for best model quality\"}\n            ],\n            \"multiSelect\": false\n        },\n        {\n            \"question\": \"Do you want to create a validation split from the training data?\",\n            \"header\": \"Split data\",\n            \"options\": [\n                {\"label\": \"Yes (Recommended)\", \"description\": \"Automatically split 15% of training data for validation\"},\n                {\"label\": \"No\", \"description\": \"Use existing validation split from dataset\"}\n            ],\n            \"multiSelect\": false\n        },\n        {\n            \"question\": \"Which GPU hardware do you want to use?\",\n            \"header\": \"Hardware Flavor\",\n            \"options\": [\n                {\"label\": \"t4-small ($0.40/hr)\", \"description\": \"1x T4, 16 GB VRAM — sufficient for all OD models under 100M params\"},\n                {\"label\": \"l4x1 ($0.80/hr)\", \"description\": \"1x L4, 24 GB VRAM — more headroom for large images or batch sizes\"},\n                {\"label\": \"a10g-large ($1.50/hr)\", \"description\": \"1x A10G, 24 GB VRAM — faster training, more CPU/RAM\"},\n                {\"label\": \"a100-large ($2.50/hr)\", \"description\": \"1x A100, 80 GB VRAM — fastest, for very large datasets or image sizes\"}\n            ],\n            \"multiSelect\": false\n        }\n    ]\n})\n```\n\n**Step 4: Prepare training script**\n\nFor object detection, use [scripts/object_detection_training.py](scripts/object_detection_training.py) as the production-ready template. For image classification, use [scripts/image_classification_training.py](scripts/image_classification_training.py). For SAM/SAM2 segmentation, use [scripts/sam_segmentation_training.py](scripts/sam_segmentation_training.py). All scripts use `HfArgumentParser` — all configuration is passed via CLI arguments in `script_args`, NOT by editing Python variables. For timm model details, see [references/timm_trainer.md](references/timm_trainer.md). For SAM2 training details, see [references/finetune_sam2_trainer.md](references/finetune_sam2_trainer.md).\n\n**Step 5: Save script, submit job, and report**\n\n1. **Save the script locally** to `submitted_jobs/` in the workspace root (create if needed) with a descriptive name like `training_<dataset>_<YYYYMMDD_HHMMSS>.py`. Tell the user the path.\n2. **Submit** using `hf_jobs` MCP tool (preferred) or `HfApi().run_uv_job()` — see directive #1 for both methods. Pass all config via `script_args`.\n3. **Report** the job ID (from `.id` attribute), monitoring URL, Trackio dashboard (`https://huggingface.co/spaces/{username}/trackio`), expected time, and estimated cost.\n4. **Wait for user** to request status checks — don't poll automatically. Training jobs run asynchronously and can take hours.\n\n## Critical directives\n\nThese rules prevent common failures. Follow them exactly.\n\n### 1. Job submission: `hf_jobs` MCP tool vs Python API\n\n**`hf_jobs()` is an MCP tool, NOT a Python function.** Do NOT try to import it from `huggingface_hub`. Call it as a tool:\n\n```\nhf_jobs(\"uv\", {\"script\": training_script_content, \"flavor\": \"a10g-large\", \"timeout\": \"4h\", \"secrets\": {\"HF_TOKEN\": \"$HF_TOKEN\"}})\n```\n\n**If `hf_jobs` MCP tool is unavailable**, use the Python API directly:\n\n```python\nfrom huggingface_hub import HfApi, get_token\napi = HfApi()\njob_info = api.run_uv_job(\n    script=\"path/to/training_script.py\",  # file PATH, NOT content\n    script_args=[\"--dataset_name\", \"cppe-5\", ...],\n    flavor=\"a10g-large\",\n    timeout=14400,  # seconds (4 hours)\n    env={\"PYTHONUNBUFFERED\": \"1\"},\n    secrets={\"HF_TOKEN\": get_token()},  # MUST use get_token(), NOT \"$HF_TOKEN\"\n)\nprint(f\"Job ID: {job_info.id}\")\n```\n\n**Critical differences between the two methods:**\n\n| | `hf_jobs` MCP tool | `HfApi().run_uv_job()` |\n|---|---|---|\n| `script` param | Python code string or URL (NOT local paths) | File path to `.py` file (NOT content) |\n| Token in secrets | `\"$HF_TOKEN\"` (auto-replaced) | `get_token()` (actual token value) |\n| Timeout format | String (`\"4h\"`) | Seconds (`14400`) |\n\n**Rules for both methods:**\n- The training script MUST include PEP 723 inline metadata with dependencies\n- Do NOT use `image` or `command` parameters (those belong to `run_job()`, not `run_uv_job()`)\n\n### 2. Authentication via job secrets + explicit hub_token injection\n\n**Job config** MUST include the token in secrets — syntax depends on submission method (see table above).\n\n**Training script requirement:** The Transformers `Trainer` calls `create_repo(token=self.args.hub_token)` during `__init__()` when `push_to_hub=True`. The training script MUST inject `HF_TOKEN` into `training_args.hub_token` AFTER parsing args but BEFORE creating the `Trainer`. The template `scripts/object_detection_training.py` already includes this:\n\n```python\nhf_token = os.environ.get(\"HF_TOKEN\")\nif training_args.push_to_hub and not training_args.hub_token:\n    if hf_token:\n        training_args.hub_token = hf_token\n```\n\nIf you write a custom script, you MUST include this token injection before the `Trainer(...)` call.\n\n- Do NOT call `login()` in custom scripts unless replicating the full pattern from `scripts/object_detection_training.py`\n- Do NOT rely on implicit token resolution (`hub_token=None`) — unreliable in Jobs\n- See the `hugging-face-jobs` skill → *Token Usage Guide* for full details\n\n### 3. JobInfo attribute\n\nAccess the job identifier using `.id` (NOT `.job_id` or `.name` — these don't exist):\n\n```python\njob_info = api.run_uv_job(...)  # or hf_jobs(\"uv\", {...})\njob_id = job_info.id  # Correct -- returns string like \"687fb701029421ae5549d998\"\n```\n\n### 4. Required training flags and HfArgumentParser boolean syntax\n\n`scripts/object_detection_training.py` uses `HfArgumentParser` — all config is passed via `script_args`. Boolean arguments have two syntaxes:\n\n- **`bool` fields** (e.g., `push_to_hub`, `do_train`): Use as bare flags (`--push_to_hub`) or negate with `--no_` prefix (`--no_remove_unused_columns`)\n- **`Optional[bool]` fields** (e.g., `greater_is_better`): MUST pass explicit value (`--greater_is_better True`). Bare `--greater_is_better` causes `error: expected one argument`\n\nRequired flags for object detection:\n\n```\n--no_remove_unused_columns          # MUST: preserves image column for pixel_values\n--no_eval_do_concat_batches         # MUST: images have different numbers of target boxes\n--push_to_hub                       # MUST: environment is ephemeral\n--hub_model_id username/model-name\n--metric_for_best_model eval_map\n--greater_is_better True            # MUST pass \"True\" explicitly (Optional[bool])\n--do_train\n--do_eval\n```\n\nRequired flags for image classification:\n\n```\n--no_remove_unused_columns          # MUST: preserves image column for pixel_values\n--push_to_hub                       # MUST: environment is ephemeral\n--hub_model_id username/model-name\n--metric_for_best_model eval_accuracy\n--greater_is_better True            # MUST pass \"True\" explicitly (Optional[bool])\n--do_train\n--do_eval\n```\n\nRequired flags for SAM/SAM2 segmentation:\n\n```\n--remove_unused_columns False       # MUST: preserves input_boxes/input_points\n--push_to_hub                       # MUST: environment is ephemeral\n--hub_model_id username/model-name\n--do_train\n--prompt_type bbox                  # or \"point\"\n--dataloader_pin_memory False       # MUST: avoids pin_memory issues with custom collator\n```\n\n### 5. Timeout management\n\nDefault 30 min is TOO SHORT for object detection. Set minimum 2-4 hours. Add 30% buffer for model loading, preprocessing, and Hub push.\n\n| Scenario | Timeout |\n|----------|---------|\n| Quick test (100-200 images, 5-10 epochs) | 1h |\n| Development (500-1K images, 15-20 epochs) | 2-3h |\n| Production (1K-5K images, 30 epochs) | 4-6h |\n| Large dataset (5K+ images) | 6-12h |\n\n### 6. Trackio monitoring\n\nTrackio is **always enabled** in the object detection training script — it calls `trackio.init()` and `trackio.finish()` automatically. No need to pass `--report_to trackio`. The project name is taken from `--output_dir` and the run name from `--run_name`. For image classification, pass `--report_to trackio` in `TrainingArguments`.\n\nDashboard at: `https://huggingface.co/spaces/{username}/trackio`\n\n## Model & hardware selection\n\n### Recommended object detection models\n\n| Model | Params | Use case |\n|-------|--------|----------|\n| `ustc-community/dfine-small-coco` | 10.4M | Best starting point — fast, cheap, SOTA quality |\n| `PekingU/rtdetr_v2_r18vd` | 20.2M | Lightweight real-time detector |\n| `ustc-community/dfine-large-coco` | 31.4M | Higher accuracy, still efficient |\n| `PekingU/rtdetr_v2_r50vd` | 43M | Strong real-time baseline |\n| `ustc-community/dfine-xlarge-obj365` | 63.5M | Best accuracy (pretrained on Objects365) |\n| `PekingU/rtdetr_v2_r101vd` | 76M | Largest RT-DETR v2 variant |\n\nStart with `ustc-community/dfine-small-coco` for fast iteration. Move to D-FINE Large or RT-DETR v2 R50 for better accuracy.\n\n### Recommended image classification models\n\nAll `timm/` models work out of the box via `AutoModelForImageClassification` (loaded as `TimmWrapperForImageClassification`). See [references/timm_trainer.md](references/timm_trainer.md) for details.\n\n| Model | Params | Use case |\n|-------|--------|----------|\n| `timm/mobilenetv3_small_100.lamb_in1k` | 2.5M | Ultra-lightweight — mobile/edge, fastest training |\n| `timm/mobilevit_s.cvnets_in1k` | 5.6M | Mobile transformer — good accuracy/speed trade-off |\n| `timm/resnet50.a1_in1k` | 25.6M | Strong CNN baseline — reliable, well-studied |\n| `timm/vit_base_patch16_dinov3.lvd1689m` | 86.6M | Best accuracy — DINOv3 self-supervised ViT |\n\nStart with `timm/mobilenetv3_small_100.lamb_in1k` for fast iteration. Move to `timm/resnet50.a1_in1k` or `timm/vit_base_patch16_dinov3.lvd1689m` for better accuracy.\n\n### Recommended SAM/SAM2 segmentation models\n\n| Model | Params | Use case |\n|-------|--------|----------|\n| `facebook/sam2.1-hiera-tiny` | 38.9M | Fastest SAM2 — good for quick experiments |\n| `facebook/sam2.1-hiera-small` | 46.0M | Best starting point — good quality/speed balance |\n| `facebook/sam2.1-hiera-base-plus` | 80.8M | Higher capacity for complex segmentation |\n| `facebook/sam2.1-hiera-large` | 224.4M | Best SAM2 accuracy — requires more VRAM |\n| `facebook/sam-vit-base` | 93.7M | Original SAM — ViT-B backbone |\n| `facebook/sam-vit-large` | 312.3M | Original SAM — ViT-L backbone |\n| `facebook/sam-vit-huge` | 641.1M | Original SAM — ViT-H, best SAM v1 accuracy |\n\nStart with `facebook/sam2.1-hiera-small` for fast iteration. SAM2 models are generally more efficient than SAM v1 at similar quality. Only the mask decoder is trained by default (vision and prompt encoders are frozen).\n\n### Hardware recommendation\n\nAll recommended OD and IC models are under 100M params — **`t4-small` (16 GB VRAM, $0.40/hr) is sufficient for all of them.** Image classification models are generally smaller and faster than object detection models — `t4-small` handles even ViT-Base comfortably. For SAM2 models up to `hiera-base-plus`, `t4-small` is sufficient since only the mask decoder is trained. For `sam2.1-hiera-large` or SAM v1 models, use `l4x1` or `a10g-large`. Only upgrade if you hit OOM from large batch sizes — reduce batch size first before switching hardware. Common upgrade path: `t4-small` → `l4x1` ($0.80/hr, 24 GB) → `a10g-large` ($1.50/hr, 24 GB).\n\nFor full hardware flavor list: refer to the `hugging-face-jobs` skill. For cost estimation: run `scripts/estimate_cost.py`.\n\n## Quick start — Object Detection\n\nThe `script_args` below are the same for both submission methods. See directive #1 for the critical differences between them.\n\n```python\nOD_SCRIPT_ARGS = [\n    \"--model_name_or_path\", \"ustc-community/dfine-small-coco\",\n    \"--dataset_name\", \"cppe-5\",\n    \"--image_square_size\", \"640\",\n    \"--output_dir\", \"dfine_finetuned\",\n    \"--num_train_epochs\", \"30\",\n    \"--per_device_train_batch_size\", \"8\",\n    \"--learning_rate\", \"5e-5\",\n    \"--eval_strategy\", \"epoch\",\n    \"--save_strategy\", \"epoch\",\n    \"--save_total_limit\", \"2\",\n    \"--load_best_model_at_end\",\n    \"--metric_for_best_model\", \"eval_map\",\n    \"--greater_is_better\", \"True\",\n    \"--no_remove_unused_columns\",\n    \"--no_eval_do_concat_batches\",\n    \"--push_to_hub\",\n    \"--hub_model_id\", \"username/model-name\",\n    \"--do_train\",\n    \"--do_eval\",\n]\n```\n\n```python\nfrom huggingface_hub import HfApi, get_token\napi = HfApi()\njob_info = api.run_uv_job(\n    script=\"scripts/object_detection_training.py\",\n    script_args=OD_SCRIPT_ARGS,\n    flavor=\"t4-small\",\n    timeout=14400,\n    env={\"PYTHONUNBUFFERED\": \"1\"},\n    secrets={\"HF_TOKEN\": get_token()},\n)\nprint(f\"Job ID: {job_info.id}\")\n```\n\n### Key OD `script_args`\n\n- `--model_name_or_path` — recommended: `\"ustc-community/dfine-small-coco\"` (see model table above)\n- `--dataset_name` — the Hub dataset ID\n- `--image_square_size` — 480 (fast iteration) or 800 (better accuracy)\n- `--hub_model_id` — `\"username/model-name\"` for Hub persistence\n- `--num_train_epochs` — 30 typical for convergence\n- `--train_val_split` — fraction to split for validation (default 0.15), set if dataset lacks a validation split\n- `--max_train_samples` — truncate training set (useful for quick test runs, e.g. `\"785\"` for ~10% of a 7.8K dataset)\n- `--max_eval_samples` — truncate evaluation set\n\n## Quick start — Image Classification\n\n```python\nIC_SCRIPT_ARGS = [\n    \"--model_name_or_path\", \"timm/mobilenetv3_small_100.lamb_in1k\",\n    \"--dataset_name\", \"ethz/food101\",\n    \"--output_dir\", \"food101_classifier\",\n    \"--num_train_epochs\", \"5\",\n    \"--per_device_train_batch_size\", \"32\",\n    \"--per_device_eval_batch_size\", \"32\",\n    \"--learning_rate\", \"5e-5\",\n    \"--eval_strategy\", \"epoch\",\n    \"--save_strategy\", \"epoch\",\n    \"--save_total_limit\", \"2\",\n    \"--load_best_model_at_end\",\n    \"--metric_for_best_model\", \"eval_accuracy\",\n    \"--greater_is_better\", \"True\",\n    \"--no_remove_unused_columns\",\n    \"--push_to_hub\",\n    \"--hub_model_id\", \"username/food101-classifier\",\n    \"--do_train\",\n    \"--do_eval\",\n]\n```\n\n```python\nfrom huggingface_hub import HfApi, get_token\napi = HfApi()\njob_info = api.run_uv_job(\n    script=\"scripts/image_classification_training.py\",\n    script_args=IC_SCRIPT_ARGS,\n    flavor=\"t4-small\",\n    timeout=7200,\n    env={\"PYTHONUNBUFFERED\": \"1\"},\n    secrets={\"HF_TOKEN\": get_token()},\n)\nprint(f\"Job ID: {job_info.id}\")\n```\n\n### Key IC `script_args`\n\n- `--model_name_or_path` — any `timm/` model or Transformers classification model (see model table above)\n- `--dataset_name` — the Hub dataset ID\n- `--image_column_name` — column containing PIL images (default: `\"image\"`)\n- `--label_column_name` — column containing class labels (default: `\"label\"`)\n- `--hub_model_id` — `\"username/model-name\"` for Hub persistence\n- `--num_train_epochs` — 3-5 typical for classification (fewer than OD)\n- `--per_device_train_batch_size` — 16-64 (classification models use less memory than OD)\n- `--train_val_split` — fraction to split for validation (default 0.15), set if dataset lacks a validation split\n- `--max_train_samples` / `--max_eval_samples` — truncate for quick tests\n\n## Quick start — SAM/SAM2 Segmentation\n\n```python\nSAM_SCRIPT_ARGS = [\n    \"--model_name_or_path\", \"facebook/sam2.1-hiera-small\",\n    \"--dataset_name\", \"merve/MicroMat-mini\",\n    \"--prompt_type\", \"bbox\",\n    \"--prompt_column_name\", \"prompt\",\n    \"--output_dir\", \"sam2-finetuned\",\n    \"--num_train_epochs\", \"30\",\n    \"--per_device_train_batch_size\", \"4\",\n    \"--learning_rate\", \"1e-5\",\n    \"--logging_steps\", \"1\",\n    \"--save_strategy\", \"epoch\",\n    \"--save_total_limit\", \"2\",\n    \"--remove_unused_columns\", \"False\",\n    \"--dataloader_pin_memory\", \"False\",\n    \"--push_to_hub\",\n    \"--hub_model_id\", \"username/sam2-finetuned\",\n    \"--do_train\",\n    \"--report_to\", \"trackio\",\n]\n```\n\n```python\nfrom huggingface_hub import HfApi, get_token\napi = HfApi()\njob_info = api.run_uv_job(\n    script=\"scripts/sam_segmentation_training.py\",\n    script_args=SAM_SCRIPT_ARGS,\n    flavor=\"t4-small\",\n    timeout=7200,\n    env={\"PYTHONUNBUFFERED\": \"1\"},\n    secrets={\"HF_TOKEN\": get_token()},\n)\nprint(f\"Job ID: {job_info.id}\")\n```\n\n### Key SAM `script_args`\n\n- `--model_name_or_path` — SAM or SAM2 model (see model table above); auto-detects SAM vs SAM2\n- `--dataset_name` — the Hub dataset ID (e.g., `\"merve/MicroMat-mini\"`)\n- `--prompt_type` — `\"bbox\"` or `\"point\"` — type of prompt in the dataset\n- `--prompt_column_name` — column with JSON-encoded prompts (default: `\"prompt\"`)\n- `--bbox_column_name` — dedicated bbox column (alternative to JSON prompt column)\n- `--point_column_name` — dedicated point column (alternative to JSON prompt column)\n- `--mask_column_name` — column with ground-truth masks (default: `\"mask\"`)\n- `--hub_model_id` — `\"username/model-name\"` for Hub persistence\n- `--num_train_epochs` — 20-30 typical for SAM fine-tuning\n- `--per_device_train_batch_size` — 2-4 (SAM models use significant memory)\n- `--freeze_vision_encoder` / `--freeze_prompt_encoder` — freeze encoder weights (default: both frozen, only mask decoder trains)\n- `--train_val_split` — fraction to split for validation (default 0.1)\n\n## Checking job status\n\n**MCP tool (if available):**\n```\nhf_jobs(\"ps\")                                   # List all jobs\nhf_jobs(\"logs\", {\"job_id\": \"your-job-id\"})      # View logs\nhf_jobs(\"inspect\", {\"job_id\": \"your-job-id\"})   # Job details\n```\n\n**Python API fallback:**\n```python\nfrom huggingface_hub import HfApi\napi = HfApi()\napi.list_jobs()                                  # List all jobs\napi.get_job_logs(job_id=\"your-job-id\")           # View logs\napi.get_job(job_id=\"your-job-id\")                # Job details\n```\n\n## Common failure modes\n\n### OOM (CUDA out of memory)\nReduce `per_device_train_batch_size` (try 4, then 2), reduce `IMAGE_SIZE`, or upgrade hardware.\n\n### Dataset format errors\nRun `scripts/dataset_inspector.py` first. The training script auto-detects xyxy vs xywh, converts string categories to integer IDs, and adds `image_id` if missing. Ensure `objects.bbox` contains 4-value coordinate lists in absolute pixels and `objects.category` contains either integer IDs or string labels.\n\n### Hub push failures (401)\nVerify: (1) job secrets include token (see directive #2), (2) script sets `training_args.hub_token` BEFORE creating the `Trainer`, (3) `push_to_hub=True` is set, (4) correct `hub_model_id`, (5) token has write permissions.\n\n### Job timeout\nIncrease timeout (see directive #5 table), reduce epochs/dataset, or use checkpoint strategy with `hub_strategy=\"every_save\"`.\n\n### KeyError: 'test' (missing test split)\nThe object detection training script handles this gracefully — it falls back to the `validation` split. Ensure you're using the latest `scripts/object_detection_training.py`.\n\n### Single-class dataset: \"iteration over a 0-d tensor\"\n`torchmetrics.MeanAveragePrecision` returns scalar (0-d) tensors for per-class metrics when there's only one class. The template `scripts/object_detection_training.py` handles this by calling `.unsqueeze(0)` on these tensors. Ensure you're using the latest template.\n\n### Poor detection performance (mAP < 0.15)\nIncrease epochs (30-50), ensure 500+ images, check per-class mAP for imbalanced classes, try different learning rates (1e-5 to 1e-4), increase image size.\n\nFor comprehensive troubleshooting: see [references/reliability_principles.md](references/reliability_principles.md)\n\n## Reference files\n\n- [scripts/object_detection_training.py](scripts/object_detection_training.py) — Production-ready object detection training script\n- [scripts/image_classification_training.py](scripts/image_classification_training.py) — Production-ready image classification training script (supports timm models)\n- [scripts/sam_segmentation_training.py](scripts/sam_segmentation_training.py) — Production-ready SAM/SAM2 segmentation training script (bbox & point prompts)\n- [scripts/dataset_inspector.py](scripts/dataset_inspector.py) — Validate dataset format for OD, classification, and SAM segmentation\n- [scripts/estimate_cost.py](scripts/estimate_cost.py) — Estimate training costs for any vision model (includes SAM/SAM2)\n- [references/object_detection_training_notebook.md](references/object_detection_training_notebook.md) — Object detection training workflow, augmentation strategies, and training patterns\n- [references/image_classification_training_notebook.md](references/image_classification_training_notebook.md) — Image classification training workflow with ViT, preprocessing, and evaluation\n- [references/finetune_sam2_trainer.md](references/finetune_sam2_trainer.md) — SAM2 fine-tuning walkthrough with MicroMat dataset, DiceCE loss, and Trainer integration\n- [references/timm_trainer.md](references/timm_trainer.md) — Using timm models with HF Trainer (TimmWrapper, transforms, full example)\n- [references/hub_saving.md](references/hub_saving.md) — Detailed Hub persistence guide and verification checklist\n- [references/reliability_principles.md](references/reliability_principles.md) — Failure prevention principles from production experience\n\n## External links\n\n- [Transformers Object Detection Guide](https://huggingface.co/docs/transformers/tasks/object_detection)\n- [Transformers Image Classification Guide](https://huggingface.co/docs/transformers/tasks/image_classification)\n- [DETR Model Documentation](https://huggingface.co/docs/transformers/model_doc/detr)\n- [ViT Model Documentation](https://huggingface.co/docs/transformers/model_doc/vit)\n- [HF Jobs Guide](https://huggingface.co/docs/huggingface_hub/guides/jobs) — Main Jobs documentation\n- [HF Jobs Configuration](https://huggingface.co/docs/hub/en/jobs-configuration) — Hardware, secrets, timeouts, namespaces\n- [HF Jobs CLI Reference](https://huggingface.co/docs/huggingface_hub/guides/cli#hf-jobs) — Command line interface\n- [Object Detection Models](https://huggingface.co/models?pipeline_tag=object-detection)\n- [Image Classification Models](https://huggingface.co/models?pipeline_tag=image-classification)\n- [SAM2 Model Documentation](https://huggingface.co/docs/transformers/model_doc/sam2)\n- [SAM Model Documentation](https://huggingface.co/docs/transformers/model_doc/sam)\n- [Object Detection Datasets](https://huggingface.co/datasets?task_categories=task_categories:object-detection)\n- [Image Classification Datasets](https://huggingface.co/datasets?task_categories=task_categories:image-classification)\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}