skill-spec turns "writing a skill" from a lone SKILL.md into a complete, verifiable, safety-bounded package — discoverable, installable, evaluable, and safely evolvable.
agents/openai.yaml exposes key-consistent metadata so an agent runtime can find and route to the skill.
Deep rules are pulled in on demand from references/, so a package copied on its own still works.
Static validation, model run, and training are reported separately — never impersonating each other.
Optimization only promotes a candidate on non-degraded held-out plus human approval. No silent overwrites.
A Codex Skill is a self-contained capability bundle an AI coding agent loads on demand. Under this standard, a skill becomes a package: a lightweight activation entry, a complete execution spec, machine-readable discovery metadata, a pinned regression contract, and an executable safety validator.
skill-spec/
├── SKILL.md # activation entry: routing / hard constraints
├── prompts/<name>.md # full execution spec: input, judgement, rules
├── agents/openai.yaml # discovery metadata (key matches directory)
├── evals/eval.yaml + 3 cases # regression contract: happy / missing / scope
├── references/ # deep rules loaded on demand
├── scripts/
│ ├── validate_skill_package.py # structural + safety + link validation
│ ├── skill_up.py # safe Skill-up adapter
│ └── skillopt.py # safe SkillOpt adapter
└── tests/test_skill_spec.py # local unit tests
The validator checks required artifacts, name/key consistency, required headings, and scans every file for credentials and absolute local paths.
Installable via pip from the release wheel — every command defaults to offline, static, safe behavior. Model-consuming runs require an explicit flag.
skill-spec-validate . # PASS: skill-spec package contract
skill-spec-skill-up validate . # validate eval schema (no model)
skill-spec-skill-up run . --execute # real model evaluation (gated)
skill-spec-skillopt preflight . # SKILLOPT_READY only when complete
python3 -m unittest discover -s tests
python3 scripts/skill_up.py validate <skill-dir>
python3 scripts/skillopt.py preflight <skill-dir>
The standard reports each layer of evidence separately and never lets one impersonate another.
Full package: entry + prompt + metadata + three eval cases + validator.
Offline structural, safety, and link checks until PASS.
skill-up validate proves the eval schema; a gated run observes real model behavior.
SkillOpt trains a candidate only; promotion needs non-degraded held-out plus a human gate.
Read the full standard in English or 中文, grab the release artifacts, or join the conversation.