← gstack WorkflowsCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to gstack Workflows
Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Compare model performance on the same bounded workflow with explicit scoring criteria.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 289
}
],
"name": "benchmark-models",
"skill_md_contents": "---\nname: benchmark-models\ndescription: Compare model performance on the same bounded workflow with explicit scoring criteria.\n---\n\n# Benchmark Models\n\nPortable ChatGPT/Codex adaptation of the `benchmark-models` workflow from `garrytan/gstack`. Preserve the original job and safety intent while mapping execution to capabilities the current host actually exposes.\n\n## When to use\n\nUse this Skill when the user explicitly names `benchmark-models` or asks for the same job described above.\n\n## Host contract\n\n- Inspect repository or file evidence before making claims about the current state.\n- Use host-native read, list, search, grep, patch, write, shell, browser, computer, and Python capabilities only when they actually exist.\n- Never claim commands, tests, browser actions, device actions, file writes, Git operations, or external mutations that were not executed.\n- Prefer read-only discovery before mutation.\n- Respect repository instructions and preserve unrelated work.\n- When the original native gstack runtime is available in Codex, it may be used as an implementation detail after inspecting the installed upstream Skill. Do not hard-code Claude-only paths as a requirement.\n\n## Workflow\n\n1. Define the metric, workload, environment, and baseline before measuring.\n2. Run the same bounded workload for each comparison target.\n3. Record execution conditions and raw evidence.\n4. Compare results without hiding variance or failed runs.\n5. Do not claim benchmark numbers unless they were actually measured.\n\n## Completion\n\nReturn the decision, findings, changed artifacts if any, executed verification, skipped checks, and remaining blockers.\n"
}SHA-256 of public snapshot: 22ef94b1c30af7e157b76f8ddd0da26f04136669841f178e1f9455df6aa6cbeb