MARKET RESEARCH

Developer Tools Codex plugins

Codex Plugin Stats: discover plugins, explore their skills and track catalog growth and observed changes.

Explore 5,308 plugins across 14 categories.

Developer Tools plugins 14

1–14 of 14

All query words must match. Use quotes for an exact phrase. Matches include publisher text, keywords and attributed research.

Discover Codex plugins by task and category
Plugin / developerCategoryFollow
CRIADOR / ReefRONALD FERRARI SOARESGuidance and references for building continual self-improvement workflows for agents with Reef.

Publisher capabilities · listing

Agent evaluation Self-improvement workflow design

Show in context →
2 more matching sources

Publisher keywords · listing

agents continual-learning self-improvement reef harness evaluation

Show in context →

Publisher full description

Guides structured evaluation, reflection, and iteration for improving agent behavior, prompts, workflows, and reliability.

Show in context →
Developer ToolsVersion 1.0.0
View details →
EvalDossierMiguel HerreroOffline verification and fixed synthetic conformance for portable signed EvalDossier dossiers.

Publisher capabilities · listing

Check signed evaluation dossier integrity offline Report the declared evidentiary basis of a verified dossier Run the fixed synthetic conformance workflow

Show in context →
2 more matching sources

Publisher keywords · listing

ai-agents evaluation attestations offline-verification

Show in context →

Publisher full description

EvalDossier checks incoming signed evaluation dossiers locally before you rely on them. It checks schema conformance, integrity, signatures, and caller-supplied audience and nonce pins, then reports the result’s declared evidentiary ba…

Show in context →
Developer ToolsVersion 0.2.1
View details →
LLM & Agent Builder CopilotKrishna SathvikDesign reliable LLM, tool, MCP, and agent workflows.

Publisher capabilities · listing

…d guardrails, human approvals, least-privilege controls, and prompt-injection defenses Create agent evaluations for task success, tool use, trajectories, side effects, latency, and cost Design tracing and observability for model calls, tools, handoffs, approvals, retries, and outcomes Use current of…

Show in context →
Developer ToolsVersion 0.1.0
View details →
Matt Skills CuratedMamdouh AboammarCurated engineering, AI/ML, and agentic productivity Skills adapted for ChatGPT and Codex.

Publisher capabilities · listing

…ry-source technical research AI & ML model engineering Self-healing data remediation ML statistical evaluation & metrics Cognitive workspace reasoning (J-space) Autonomous goal contracts Skill conductor authoring & evals Git safety guardrails & pre-commit Workflow design & session handoffs Course sc…

Show in context →
Developer ToolsVersion 1.1.0
View details →
CompText BenchmarkCompText LabsRun, inspect, and compare reproducible Raw vs CompText benchmarks with quality-first evidence.

Publisher keywords · listing

benchmark context-compression evaluation comptext

Show in context →
Developer ToolsVersion 0.1.5
View details →
Plugin EvalOpenAIEvaluate Codex skills and plugins from chat with a beginner-friendly start command, local-first reports, token budget explanations, and g...

Publisher keywords · package

codex plugin skill evaluation quality budget

Show in context →
Developer ToolsVersion 0.1.2
View details →
RAG & GenAI CopilotKrishna SathvikProduction-focused RAG architecture, debugging, evaluation, security, and reliability.

Publisher keywords · listing

rag genai retrieval llm evaluation security observability

Show in context →
2 more matching sources

Publisher description

Production-focused RAG architecture, debugging, evaluation, security, and reliability.

Show in context →

Publisher full description

…ng, indexing, retrieval, reranking, context construction, generation, citations, authorization, and evaluation; design practical RAG architectures; build retrieval and groundedness evals; review prompt-injection and data-leakage risks; and improve observability, latency, and cost. It favors measurab…

Show in context →
Developer ToolsVersion 0.1.0
View details →
Brainbase MCPBrainbase Labs

Publisher description

…T and Codex. It supports revision-safe agent changes, registry skills and MCP server configuration, evaluations, orchestrations, schedules, and task runs, with explicit confirmation for destructive actions and OAuth-based access to the user's Brainbase workspace.

Show in context →
Developer ToolsVersion 1.0.0
View details →
CovalCoval

Publisher description

Coval helps teams inspect AI agents, test sets, personas, metrics, and evaluation runs; create and refine evaluation resources; launch evaluations; and ask Sofia for read-only analysis grounded in their Coval organization.

Show in context →
Developer ToolsVersion 2.0.0
View details →
get-fableMamdouh Abo Ammar

Publisher description

…search, planning, test-first changes, delegation, verification, review, security, release, handoff, evaluation, and recovery across 25 canonical skills.

Show in context →
Developer ToolsVersion 1.5.1
View details →
Git Diff Patcher BridgeMuhammet Avcı

Publisher description

…boundaries. The connector also supports a seeded Demo Workspace for zero-install reviewer and user evaluation.

Show in context →
Developer ToolsVersion 2.0.0
View details →
Intuitive Software DesignArcanEdge LLC

Publisher full description

…s focused guidance, formal audits, smallest-complete-change improvements, and an 11-case behavioral evaluation kit.

Show in context →
Developer ToolsVersion 1.3.1
View details →
OpenlayerOpenlayer

Publisher description

Openlayer helps teams inspect AI projects, data sources, traces, tests, evaluation results, and governance controls, and manage private workspace resources through ChatGPT.

Show in context →
Developer ToolsVersion 1.0.0
View details →
Prompt EngineerThe Doers Firm LTD

Publisher full description

…clear, model-aware prompts; improve existing prompts while preserving intent; and design practical evaluation cases to compare quality, reliability, and safety. Distinguish prompt-only behavior from capabilities that require tools or application code. Treat prompt results as empirical, model-depend…

Show in context →
Developer ToolsVersion 0.1.0
View details →