MARKET RESEARCH

Developer Tools Codex plugins

Codex Plugin Stats: discover plugins, explore their skills and track catalog growth and observed changes.

Explore 5,295 plugins across 14 categories.

Developer Tools plugins 25

1–25 of 25

All query words must match. Use quotes for an exact phrase. Matches include publisher text, keywords and attributed research.

Discover Codex plugins by task and category
Plugin / developerCategoryFollow
EvalDossierMiguel HerreroOffline verification and fixed synthetic conformance for portable signed EvalDossier dossiers.

Plugin name

EvalDossier

Show in context →
5 more matching sources

Package name

evaldossier

Show in context →

Publisher capabilities · listing

Check signed evaluation dossier integrity offline Report the declared evidentiary basis of a verified dossier Run the fixed synthetic conformance workflow

Show in context →

Publisher keywords · listing

ai-agents evaluation attestations offline-verification

Show in context →

Publisher description

Offline verification and fixed synthetic conformance for portable signed EvalDossier dossiers.

Show in context →

Publisher full description

EvalDossier checks incoming signed evaluation dossiers locally before you rely on them. It checks schema conformance, integrity, signatures, and caller-supplied audience and nonce pins, then reports t…

Show in context →
Developer ToolsVersion 0.2.1
View details →
Plugin EvalOpenAIEvaluate Codex skills and plugins from chat with a beginner-friendly start command, local-first reports, token budget explanations, and g...

Plugin name

Plugin Eval

Show in context →
5 more matching sources

Package name

plugin-eval

Show in context →

Publisher subtitle

Start from chat, then evaluate or benchmark locally

Show in context →

Publisher keywords · package

codex plugin skill evaluation quality budget

Show in context →

Publisher description

Evaluate Codex skills and plugins from chat with a beginner-friendly start command, local-first reports, token budget explanations, and guided benchmarking.

Show in context →

Publisher full description

Ask Codex to evaluate a plugin or skill, give you a full analysis of a named plugin such as game-studio, explain why it scored that way, show what to fix first, explain its token budget, measure real token usage,…

Show in context →
Developer ToolsVersion 0.1.2
View details →
CovalCovalCoval helps teams inspect AI agents, test sets, personas, metrics, and evaluation runs; create and refine evaluation resources; launch ev...

Publisher subtitle

Evaluate voice and chat agents

Show in context →
1 more matching sources

Publisher description

Coval helps teams inspect AI agents, test sets, personas, metrics, and evaluation runs; create and refine evaluation resources; launch evaluations; and ask Sofia for read-only analysis grounded in their Coval organization.

Show in context →
Developer ToolsVersion 2.0.0
View details →
Agent ReachRONALD FERRARI SOARESAgent Reach integration skill for routing internet research and read-only retrieval across supported web, social, developer, career, and ...

Publisher capabilities · listing

Research routing Read-only retrieval

Show in context →
1 more matching sources

Publisher description

Agent Reach integration skill for routing internet research and read-only retrieval across supported web, social, developer, career, and media sources.

Show in context →
Developer ToolsVersion 1.5.0
View details →
CRIADOR / ReefRONALD FERRARI SOARESGuidance and references for building continual self-improvement workflows for agents with Reef.

Publisher capabilities · listing

Agent evaluation Self-improvement workflow design

Show in context →
2 more matching sources

Publisher keywords · listing

agents continual-learning self-improvement reef harness evaluation

Show in context →

Publisher full description

Guides structured evaluation, reflection, and iteration for improving agent behavior, prompts, workflows, and reliability.

Show in context →
Developer ToolsVersion 1.0.0
View details →
LLM & Agent Builder CopilotKrishna SathvikDesign reliable LLM, tool, MCP, and agent workflows.

Publisher capabilities · listing

…s, idempotency, errors, and approval requirements Design manager, handoff, router, parallel-worker, evaluator, and orchestrator patterns Plan MCP tools, resources, prompts, authentication, long-running tasks, and observability Design conversation state, workflow state, durable preferences, and appli…

Show in context →
2 more matching sources

Publisher keywords · listing

llm agents mcp tools guardrails evals orchestration

Show in context →

Publisher full description

…n state and memory; plan handoffs and orchestration; add guardrails and human approvals; and create eval and observability strategies. It favors structured outputs, least privilege, bounded execution, recoverable state, and simple architectures.

Show in context →
Developer ToolsVersion 0.1.0
View details →
Matt Skills CuratedMamdouh AboammarCurated engineering, AI/ML, and agentic productivity Skills adapted for ChatGPT and Codex.

Publisher capabilities · listing

…ry-source technical research AI & ML model engineering Self-healing data remediation ML statistical evaluation & metrics Cognitive workspace reasoning (J-space) Autonomous goal contracts Skill conductor authoring & evals Git safety guardrails & pre-commit Workflow design & session handoffs Course sc…

Show in context →
Developer ToolsVersion 1.1.0
View details →
RAG & GenAI CopilotKrishna SathvikProduction-focused RAG architecture, debugging, evaluation, security, and reliability.

Publisher capabilities · listing

Design production RAG architectures for quality, freshness, security, latency, and cost Debug retrieval and grounded-generation failures from traces, logs, prompts, and evals Diagnose parsing, chunking, indexing, metadata, filtering, reranking, and context issues Evaluate retrieval quality, grounde…

Show in context →
3 more matching sources

Publisher keywords · listing

rag genai retrieval llm evaluation security observability

Show in context →

Publisher description

Production-focused RAG architecture, debugging, evaluation, security, and reliability.

Show in context →

Publisher full description

Production RAG & GenAI Copilot helps you design, debug, evaluate, secure, and operate retrieval-augmented generation systems. It can trace failures across ingestion, parsing, chunking, indexing, retrieval, reranking, context construction, generation, citat…

Show in context →
Developer ToolsVersion 0.1.0
View details →
ApprenticeAbhishek Samar SinghNotice a repeatable, expensive LLM call in your code and mention Apprentice: capture real examples, optimize the prompt or fine-tune a sm...

Publisher keywords · listing

llm cost-optimization prompt-optimization fine-tuning evals dspy gepa

Show in context →
1 more matching sources

Publisher description

…e real examples, optimize the prompt or fine-tune a small open model, verified by your own held-out evals before any traffic moves.

Show in context →
Developer ToolsVersion 0.1.0
View details →
CompText BenchmarkCompText LabsRun, inspect, and compare reproducible Raw vs CompText benchmarks with quality-first evidence.

Publisher keywords · listing

benchmark context-compression evaluation comptext

Show in context →
Developer ToolsVersion 0.1.5
View details →
PineconePineconePinecone vector database integration for Codex. Streamline your Pinecone development with reusable skills for managing vector indexes, qu...

Publisher keywords · listing

pinecone semantic search retrieval vector search vector database retrieval augmented generation rag agentic rag sparse search assistant pinecone assistant rag chatbot full text search bm25 hybrid search

Show in context →
Developer ToolsVersion 1.0.1
View details →
PromptfooPromptfooAgent skills for connecting, evaluating, and red teaming LLM applications with Promptfoo.

Publisher keywords · listing

eval redteam llm testing promptfoo

Show in context →
2 more matching sources

Publisher description

Agent skills for connecting, evaluating, and red teaming LLM applications with Promptfoo.

Show in context →

Publisher full description

Promptfoo skills for configuring providers and targets, writing eval suites, and setting up or running red team workflows against LLM apps.

Show in context →
Developer ToolsVersion 0.1.3
View details →
Brainbase MCPBrainbase Labs

Publisher description

…T and Codex. It supports revision-safe agent changes, registry skills and MCP server configuration, evaluations, orchestrations, schedules, and task runs, with explicit confirmation for destructive actions and OAuth-based access to the user's Brainbase workspace.

Show in context →
Developer ToolsVersion 1.0.0
View details →
Chronos for CodexDravara, LLC

Publisher full description

…un one local Governor that discovers active Codex tasks from one complete host inventory per cycle, evaluates actionable Heartbeats, and contacts only the exact affected task when intervention is justified. Optional silent hooks accelerate lifecycle hints without creating model turns. Chronos also d…

Show in context →
Developer ToolsVersion 0.9.2
View details →
GEO Tool Checktrack by track GmbH

Publisher description

…d by AI search systems. It scores page readiness, checks whether major AI crawlers are blocked, and evaluates draft passages for citability. Checks run without an account or API key and use the same scoring logic as geo-tool.com.

Show in context →
Developer ToolsVersion 1.0.0
View details →
get-fableMamdouh Abo Ammar

Publisher description

…search, planning, test-first changes, delegation, verification, review, security, release, handoff, evaluation, and recovery across 25 canonical skills.

Show in context →
Developer ToolsVersion 1.5.1
View details →
Git Diff Patcher BridgeMuhammet Avcı

Publisher description

…boundaries. The connector also supports a seeded Demo Workspace for zero-install reviewer and user evaluation.

Show in context →
Developer ToolsVersion 2.0.0
View details →
Intuitive Software DesignArcanEdge LLC

Publisher full description

…ghs, product mental models, decision support, multi-role handoffs, and connected experience design. Evaluates UI, Flow, and Feel, including relevant device, platform, service, state, and recovery boundaries. Includes focused guidance, formal audits, smallest-complete-change improvements, and an 11-c…

Show in context →
Developer ToolsVersion 1.3.1
View details →
MailChannelsMailChannels Corporation

Publisher full description

Evaluate MailChannels and implement its Email API with official JavaScript, Python, and PHP SDK guidance plus production controls for domains, webhooks, suppressions, multi-tenant operations, and safe…

Show in context →
Developer ToolsVersion 1.0.1
View details →
omgskillsomgskills

Publisher description

…skills, and explore skills from specific creators. omgskills is read-only: it helps people find and evaluate skills without installing or changing anything on their device.

Show in context →
Developer ToolsVersion 1.0.0
View details →
OpenlayerOpenlayer

Publisher description

Openlayer helps teams inspect AI projects, data sources, traces, tests, evaluation results, and governance controls, and manage private workspace resources through ChatGPT.

Show in context →
Developer ToolsVersion 1.0.0
View details →
Prompt EngineerThe Doers Firm LTD

Publisher description

Design, improve, and evaluate clear, reliable AI prompts.

Show in context →
1 more matching sources

Publisher full description

…clear, model-aware prompts; improve existing prompts while preserving intent; and design practical evaluation cases to compare quality, reliability, and safety. Distinguish prompt-only behavior from capabilities that require tools or application code. Treat prompt results as empirical, model-depend…

Show in context →
Developer ToolsVersion 0.1.0
View details →
TavilyTavily

Publisher description

…Agents. A purpose-built web API for real-time search, scraping, crawling, and structured data retrieval. AI-native enterprises trust Tavily for data enrichment, research, RAG pipelines, and autonomous AI systems-ensuring fresh, accurate, and scalable intelligence.

Show in context →
Developer ToolsVersion 2.0.0
View details →
UniformUniform Systems, Inc.

Publisher full description

…stays small. Guidance is verified against the shipped Uniform SDK packages and measured with an A/B eval harness in the public repository.

Show in context →
Developer ToolsVersion 1.0.0
View details →
Vioscale AIVioscale Technologies Ltd

Publisher description

Vioscale AI helps users discover, compare, and evaluate software using evidence-backed facts, transparent scores, and canonical citations.

Show in context →
Developer ToolsVersion 1.0.0
View details →