12 KiB
MANIFEST POLICY
This document defines how agents should maintain the Agent Skills manifest at skvlt.yaml.
It covers four outcomes:
- addition
- replacement
- rejection
- escalation to humans
The policy is intentionally conservative because the manifest is global. A bad addition creates global maintenance cost, trigger noise, and weaker default skill selection. The goal is not merely to minimize overlap, but to keep the most complete, practical, and worth-retaining skills in each category.
Audience
This policy is for:
- agents proposing changes to
skvlt.yaml - humans reviewing manifest changes
- future automation that scans, scores, and patches the manifest
Scope
This policy applies to:
- adding a skill to an existing source block
- replacing one installed skill with another
- rejecting candidate skills
- escalating ambiguous or high-impact cases to humans
- maintaining manifest integrity fields such as
total_sources,total_skills, and per-sourcecount
This policy does not define:
- how to install skills on disk
- how to evaluate runtime correctness of a skill beyond metadata, prerequisites, and trigger scope
- how to author new skills
Operating Model
Manifest maintenance should be split across small agent roles.
Scout Agent
Finds candidate additions, upstream updates, and removal opportunities from approved sources.
Metadata Agent
Extracts:
- source repository
- skill name
- trigger description
- documentation completeness signal
- practical utility signal
- depth or breadth of reusable guidance
- explicit prerequisites
- references to a specific agent, repository, or runtime
- dependency skills
- external API or paid service requirements
- installation count
- author trust signal
- maintenance activity signal
Taxonomy Agent
Maps each skill to a task category such as:
- browser automation
- documentation authoring
- Cloudflare development
- LLM security
- skill authoring
- social research
Comparator Agent
Compares a candidate against installed skills in the same category.
Its job is to answer:
- Is this a general skill or a specialized skill?
- Which skill is more complete, practical, and worth retaining if only one remains?
- Does it substantially overlap an installed skill's trigger scope?
- If there is overlap, is it "general vs. general" or "general vs. specialized"?
- Does the overlap create redundant surface area, or does it add meaningful complementary depth?
- Does it introduce new constraints that make it less portable?
Decision Agent
Returns exactly one outcome:
addreplacerejectescalate
Manifest Editor Agent
Prepares the patch to skvlt.yaml, updates counts, and records the rationale.
Human Reviewer
Approves only escalated changes or batched high-impact changes.
Core Principles
1. Conservative Global Bias
skvlt.yaml is a global manifest. The default should be to avoid broad, noisy, or brittle skills unless they add clear value. However, reducing noise is a secondary constraint after choosing the skill that is strongest to keep long-term within its category.
2. Retention Value First
When comparing skills in the same category, first ask which one is more complete, more practical, and more worth retaining for repeated real-world use.
Completeness and practical utility should outrank superficial trigger neatness. If one skill is materially better as a lasting default, that matters more than preserving a narrower trigger boundary.
3. Trigger Overlap Matters
Substantial similarity of task triggers means two skills occupy the same task category without a real capability distinction. Overlap should be judged after content quality and retention value are assessed, not before.
4. General and Specialized Can Coexist
Do not force a trade-off between a broad general skill and a clearly narrower specialized skill.
Trade-offs are required only when the comparison is effectively "general vs. general".
5. Trust and Maintenance Break Ties
When quality and practical value are otherwise close, prefer:
- official or domain-relevant authors
- repositories with strong installation counts
- active maintenance
6. Dependencies Must Close
A skill should not be auto-added if it depends on another skill that is missing from the manifest, unless the same change adds the dependency or a human approves an exception.
Decision Workflow
- The Scout Agent identifies a candidate.
- The Metadata Agent extracts the candidate's signals.
- The Taxonomy Agent places it into a task category.
- A hard exclusion pass runs first.
- The Comparator Agent evaluates overlap against installed skills in the same category.
- The Decision Agent chooses
add,replace,reject, orescalate. - If the decision is
addorreplace, the Manifest Editor Agent drafts a patch. - If the decision is
escalate, a human receives a short review packet.
Hard Exclusion Rules
The Decision Agent should immediately return reject when any of the following is true, unless the skill is explicitly whitelisted:
- the skill is tightly bound to a specific agent
- the skill assumes a specific repository context
- the skill depends on a runtime or execution model that is outside the manifest's baseline
- the skill's prerequisites cannot be satisfied from the current manifest or the same proposed patch
- the trigger description is broad but the actual capability is narrow and misleading
Action Policy
Addition
Return add only when all of the following are true:
- the candidate survives hard exclusion
- its task category is already allowed in the manifest
- it is either specialized, genuinely non-overlapping, or clearly complementary after comparing retention value against incumbents
- its dependencies close cleanly
- it does not introduce broad trigger ambiguity without enough added practical value to justify coexistence
Typical examples:
- adding a specialized companion skill under an already trusted source
- adding a dependency skill already referenced by installed skills
Replacement
Return replace only when all of the following are true:
- the candidate and incumbent are both general skills
- their trigger scopes are substantially similar enough that keeping both mostly duplicates maintenance surface area
- the candidate is clearly more complete, practical, and worth retaining for repeated use
- the candidate is at least as portable as the incumbent
- trust and maintenance signals do not materially undercut the candidate
- there is no meaningful capability loss
Replacement should not be used for "general vs. specialized" comparisons.
Rejection
Return reject when any of the following is true:
- the skill is bound to a specific agent, repository, or runtime
- it is a weaker or lower-retention-value duplicate of an installed general skill
- it creates dependency debt
- it adds noise without enough incremental practical value
- it loses on completeness, practicality, and retention value, with trust and maintenance unable to reverse that judgment
Escalation
Return escalate when the case is not safely automatable.
Escalation is required when:
- the overlap judgment depends on interpretation rather than clear capability boundaries
- the completeness or practical-utility comparison is close rather than clear
- the change adds a brand-new source block
- the change removes the last skill from a source block
- the candidate introduces external auth, paid APIs, or unusual runtime assumptions
- content quality and trust signals point in different directions
- the skill shapes broad workflow behavior across many tasks
Scoring Rubric
Use the scoring rubric only after the hard exclusion pass.
Positive Signals
+4clearly more complete and practical than the incumbent in the same category+3clear specialization with distinct capability+2official or strongly domain-relevant author+2active maintenance+2strong installation signal+2dependency closure is already satisfied+1complements an already trusted source
Negative Signals
-4shallower, less actionable, or less worth retaining than an incumbent-3vague or noisy trigger wording-3external service requirement that is not already normal for the manifest-4missing dependency-4portability concerns
Hard Stop
reject immediatelyfor explicit agent / repo / runtime binding unless whitelisted
Decision Bands
score >= 5: eligible foraddif overlap review still shows clear incremental retention valuescore 2 to 4: preferescalatescore <= 1: preferreject
For replace, compare candidate score against incumbent score. Auto-replace only when the candidate is ahead by at least 3 points, clearly ahead on completeness and practical value, and no escalation condition is present.
Human Escalation Packet
When escalation is required, provide a short packet with:
- candidate skill
- source repository
- task category
- incumbent skill, if any
- completeness and practical utility comparison
- overlap summary
- binding and dependency notes
- trust and maintenance comparison
- recommended action
- confidence level
Manifest Invariants
Every accepted patch must preserve the following:
total_sourcesmatches the number of source blockstotal_skillsmatches the sum of per-source counts- each source
countmatches the number of listed skills - no duplicate skill names appear in the manifest
- no hard-reject skill is present unless explicitly whitelisted
- dependency exceptions are documented in the review note
Change Classes
Treat these classes differently:
Safe to Automate
- adding a skill to an existing trusted source
- adding a missing dependency that is already referenced by installed skills
- replacing a lower-value general skill with a clearly more complete and practical general skill
Must Escalate
- new source added
- source removed
- process-shaping skill added, removed, or replaced
- ambiguous overlap between two broad general skills with close retention value
- any change that would alter maintenance policy itself
Recommended Cadence
Weekly
- scout for new candidates in already approved sources
- check for dependency gaps and broken references
Monthly
- review overlapping general skills
- re-evaluate trust and maintenance signals
Quarterly
- review category taxonomy
- review source allowlist
- review any existing whitelists for otherwise excluded skills
Output Format for Agents
Every maintenance run should emit a compact decision log.
Recommended fields:
{
"candidate": "nanobanana",
"source": "resciencelab/opc-skills",
"category": "image generation",
"incumbent": null,
"action": "add",
"reason": "dependency closure for installed skills",
"confidence": 0.91,
"needs_human": false
}
Current Manifest Guidance
For the current skvlt.yaml:
- broad workflow skills should be treated as high-impact
- source-level edits should escalate more readily than skill-level edits
- when two broad general skills overlap, first ask which one would still be worth keeping if only one remained
- prefer the skill that is more complete, more practical, and more worth retaining long-term; use trigger noise, authority, and maintenance as secondary constraints
- dependency-completion additions inside an existing trusted source can be auto-approved
This means a change like "add a missing dependency skill under an existing source" is lower risk than "replace a broad browser automation skill family", but for overlapping general skills the primary question should be retention value rather than trigger neatness alone.