340 lines
12 KiB
Markdown
340 lines
12 KiB
Markdown
# MANIFEST POLICY
|
|
|
|
This document defines how agents should maintain the Agent Skills manifest at `skvlt.yaml`.
|
|
|
|
It covers four outcomes:
|
|
|
|
- addition
|
|
- replacement
|
|
- rejection
|
|
- escalation to humans
|
|
|
|
The policy is intentionally conservative because the manifest is global. A bad addition creates global maintenance cost, trigger noise, and weaker default skill selection. The goal is not merely to minimize overlap, but to keep the most complete, practical, and worth-retaining skills in each category.
|
|
|
|
## Audience
|
|
|
|
This policy is for:
|
|
|
|
- agents proposing changes to `skvlt.yaml`
|
|
- humans reviewing manifest changes
|
|
- future automation that scans, scores, and patches the manifest
|
|
|
|
## Scope
|
|
|
|
This policy applies to:
|
|
|
|
- adding a skill to an existing source block
|
|
- replacing one installed skill with another
|
|
- rejecting candidate skills
|
|
- escalating ambiguous or high-impact cases to humans
|
|
- maintaining manifest integrity fields such as `total_sources`, `total_skills`, and per-source `count`
|
|
|
|
This policy does not define:
|
|
|
|
- how to install skills on disk
|
|
- how to evaluate runtime correctness of a skill beyond metadata, prerequisites, and trigger scope
|
|
- how to author new skills
|
|
|
|
## Operating Model
|
|
|
|
Manifest maintenance should be split across small agent roles.
|
|
|
|
### Scout Agent
|
|
|
|
Finds candidate additions, upstream updates, and removal opportunities from approved sources.
|
|
|
|
### Metadata Agent
|
|
|
|
Extracts:
|
|
|
|
- source repository
|
|
- skill name
|
|
- trigger description
|
|
- documentation completeness signal
|
|
- practical utility signal
|
|
- depth or breadth of reusable guidance
|
|
- explicit prerequisites
|
|
- references to a specific agent, repository, or runtime
|
|
- dependency skills
|
|
- external API or paid service requirements
|
|
- installation count
|
|
- author trust signal
|
|
- maintenance activity signal
|
|
|
|
### Taxonomy Agent
|
|
|
|
Maps each skill to a task category such as:
|
|
|
|
- browser automation
|
|
- documentation authoring
|
|
- Cloudflare development
|
|
- LLM security
|
|
- skill authoring
|
|
- social research
|
|
|
|
### Comparator Agent
|
|
|
|
Compares a candidate against installed skills in the same category.
|
|
|
|
Its job is to answer:
|
|
|
|
- Is this a general skill or a specialized skill?
|
|
- Which skill is more complete, practical, and worth retaining if only one remains?
|
|
- Does it substantially overlap an installed skill's trigger scope?
|
|
- If there is overlap, is it "general vs. general" or "general vs. specialized"?
|
|
- Does the overlap create redundant surface area, or does it add meaningful complementary depth?
|
|
- Does it introduce new constraints that make it less portable?
|
|
|
|
### Decision Agent
|
|
|
|
Returns exactly one outcome:
|
|
|
|
- `add`
|
|
- `replace`
|
|
- `reject`
|
|
- `escalate`
|
|
|
|
### Manifest Editor Agent
|
|
|
|
Prepares the patch to `skvlt.yaml`, updates counts, and records the rationale.
|
|
|
|
### Human Reviewer
|
|
|
|
Approves only escalated changes or batched high-impact changes.
|
|
|
|
## Core Principles
|
|
|
|
### 1. Conservative Global Bias
|
|
|
|
`skvlt.yaml` is a global manifest. The default should be to avoid broad, noisy, or brittle skills unless they add clear value. However, reducing noise is a secondary constraint after choosing the skill that is strongest to keep long-term within its category.
|
|
|
|
### 2. Retention Value First
|
|
|
|
When comparing skills in the same category, first ask which one is more complete, more practical, and more worth retaining for repeated real-world use.
|
|
|
|
Completeness and practical utility should outrank superficial trigger neatness. If one skill is materially better as a lasting default, that matters more than preserving a narrower trigger boundary.
|
|
|
|
### 3. Trigger Overlap Matters
|
|
|
|
Substantial similarity of task triggers means two skills occupy the same task category without a real capability distinction. Overlap should be judged after content quality and retention value are assessed, not before.
|
|
|
|
### 4. General and Specialized Can Coexist
|
|
|
|
Do not force a trade-off between a broad general skill and a clearly narrower specialized skill.
|
|
|
|
Trade-offs are required only when the comparison is effectively "general vs. general".
|
|
|
|
### 5. Trust and Maintenance Break Ties
|
|
|
|
When quality and practical value are otherwise close, prefer:
|
|
|
|
- official or domain-relevant authors
|
|
- repositories with strong installation counts
|
|
- active maintenance
|
|
|
|
### 6. Dependencies Must Close
|
|
|
|
A skill should not be auto-added if it depends on another skill that is missing from the manifest, unless the same change adds the dependency or a human approves an exception.
|
|
|
|
## Decision Workflow
|
|
|
|
1. The Scout Agent identifies a candidate.
|
|
2. The Metadata Agent extracts the candidate's signals.
|
|
3. The Taxonomy Agent places it into a task category.
|
|
4. A hard exclusion pass runs first.
|
|
5. The Comparator Agent evaluates overlap against installed skills in the same category.
|
|
6. The Decision Agent chooses `add`, `replace`, `reject`, or `escalate`.
|
|
7. If the decision is `add` or `replace`, the Manifest Editor Agent drafts a patch.
|
|
8. If the decision is `escalate`, a human receives a short review packet.
|
|
|
|
## Hard Exclusion Rules
|
|
|
|
The Decision Agent should immediately return `reject` when any of the following is true, unless the skill is explicitly whitelisted:
|
|
|
|
- the skill is tightly bound to a specific agent
|
|
- the skill assumes a specific repository context
|
|
- the skill depends on a runtime or execution model that is outside the manifest's baseline
|
|
- the skill's prerequisites cannot be satisfied from the current manifest or the same proposed patch
|
|
- the trigger description is broad but the actual capability is narrow and misleading
|
|
|
|
## Action Policy
|
|
|
|
### Addition
|
|
|
|
Return `add` only when all of the following are true:
|
|
|
|
- the candidate survives hard exclusion
|
|
- its task category is already allowed in the manifest
|
|
- it is either specialized, genuinely non-overlapping, or clearly complementary after comparing retention value against incumbents
|
|
- its dependencies close cleanly
|
|
- it does not introduce broad trigger ambiguity without enough added practical value to justify coexistence
|
|
|
|
Typical examples:
|
|
|
|
- adding a specialized companion skill under an already trusted source
|
|
- adding a dependency skill already referenced by installed skills
|
|
|
|
### Replacement
|
|
|
|
Return `replace` only when all of the following are true:
|
|
|
|
- the candidate and incumbent are both general skills
|
|
- their trigger scopes are substantially similar enough that keeping both mostly duplicates maintenance surface area
|
|
- the candidate is clearly more complete, practical, and worth retaining for repeated use
|
|
- the candidate is at least as portable as the incumbent
|
|
- trust and maintenance signals do not materially undercut the candidate
|
|
- there is no meaningful capability loss
|
|
|
|
Replacement should not be used for "general vs. specialized" comparisons.
|
|
|
|
### Rejection
|
|
|
|
Return `reject` when any of the following is true:
|
|
|
|
- the skill is bound to a specific agent, repository, or runtime
|
|
- it is a weaker or lower-retention-value duplicate of an installed general skill
|
|
- it creates dependency debt
|
|
- it adds noise without enough incremental practical value
|
|
- it loses on completeness, practicality, and retention value, with trust and maintenance unable to reverse that judgment
|
|
|
|
### Escalation
|
|
|
|
Return `escalate` when the case is not safely automatable.
|
|
|
|
Escalation is required when:
|
|
|
|
- the overlap judgment depends on interpretation rather than clear capability boundaries
|
|
- the completeness or practical-utility comparison is close rather than clear
|
|
- the change adds a brand-new source block
|
|
- the change removes the last skill from a source block
|
|
- the candidate introduces external auth, paid APIs, or unusual runtime assumptions
|
|
- content quality and trust signals point in different directions
|
|
- the skill shapes broad workflow behavior across many tasks
|
|
|
|
## Scoring Rubric
|
|
|
|
Use the scoring rubric only after the hard exclusion pass.
|
|
|
|
### Positive Signals
|
|
|
|
- `+4` clearly more complete and practical than the incumbent in the same category
|
|
- `+3` clear specialization with distinct capability
|
|
- `+2` official or strongly domain-relevant author
|
|
- `+2` active maintenance
|
|
- `+2` strong installation signal
|
|
- `+2` dependency closure is already satisfied
|
|
- `+1` complements an already trusted source
|
|
|
|
### Negative Signals
|
|
|
|
- `-4` shallower, less actionable, or less worth retaining than an incumbent
|
|
- `-3` vague or noisy trigger wording
|
|
- `-3` external service requirement that is not already normal for the manifest
|
|
- `-4` missing dependency
|
|
- `-4` portability concerns
|
|
|
|
### Hard Stop
|
|
|
|
- `reject immediately` for explicit agent / repo / runtime binding unless whitelisted
|
|
|
|
### Decision Bands
|
|
|
|
- `score >= 5`: eligible for `add` if overlap review still shows clear incremental retention value
|
|
- `score 2 to 4`: prefer `escalate`
|
|
- `score <= 1`: prefer `reject`
|
|
|
|
For `replace`, compare candidate score against incumbent score. Auto-replace only when the candidate is ahead by at least `3` points, clearly ahead on completeness and practical value, and no escalation condition is present.
|
|
|
|
## Human Escalation Packet
|
|
|
|
When escalation is required, provide a short packet with:
|
|
|
|
- candidate skill
|
|
- source repository
|
|
- task category
|
|
- incumbent skill, if any
|
|
- completeness and practical utility comparison
|
|
- overlap summary
|
|
- binding and dependency notes
|
|
- trust and maintenance comparison
|
|
- recommended action
|
|
- confidence level
|
|
|
|
## Manifest Invariants
|
|
|
|
Every accepted patch must preserve the following:
|
|
|
|
- `total_sources` matches the number of source blocks
|
|
- `total_skills` matches the sum of per-source counts
|
|
- each source `count` matches the number of listed skills
|
|
- no duplicate skill names appear in the manifest
|
|
- no hard-reject skill is present unless explicitly whitelisted
|
|
- dependency exceptions are documented in the review note
|
|
|
|
## Change Classes
|
|
|
|
Treat these classes differently:
|
|
|
|
### Safe to Automate
|
|
|
|
- adding a skill to an existing trusted source
|
|
- adding a missing dependency that is already referenced by installed skills
|
|
- replacing a lower-value general skill with a clearly more complete and practical general skill
|
|
|
|
### Must Escalate
|
|
|
|
- new source added
|
|
- source removed
|
|
- process-shaping skill added, removed, or replaced
|
|
- ambiguous overlap between two broad general skills with close retention value
|
|
- any change that would alter maintenance policy itself
|
|
|
|
## Recommended Cadence
|
|
|
|
### Weekly
|
|
|
|
- scout for new candidates in already approved sources
|
|
- check for dependency gaps and broken references
|
|
|
|
### Monthly
|
|
|
|
- review overlapping general skills
|
|
- re-evaluate trust and maintenance signals
|
|
|
|
### Quarterly
|
|
|
|
- review category taxonomy
|
|
- review source allowlist
|
|
- review any existing whitelists for otherwise excluded skills
|
|
|
|
## Output Format for Agents
|
|
|
|
Every maintenance run should emit a compact decision log.
|
|
|
|
Recommended fields:
|
|
|
|
```json
|
|
{
|
|
"candidate": "nanobanana",
|
|
"source": "resciencelab/opc-skills",
|
|
"category": "image generation",
|
|
"incumbent": null,
|
|
"action": "add",
|
|
"reason": "dependency closure for installed skills",
|
|
"confidence": 0.91,
|
|
"needs_human": false
|
|
}
|
|
```
|
|
|
|
## Current Manifest Guidance
|
|
|
|
For the current `skvlt.yaml`:
|
|
|
|
- broad workflow skills should be treated as high-impact
|
|
- source-level edits should escalate more readily than skill-level edits
|
|
- when two broad general skills overlap, first ask which one would still be worth keeping if only one remained
|
|
- prefer the skill that is more complete, more practical, and more worth retaining long-term; use trigger noise, authority, and maintenance as secondary constraints
|
|
- dependency-completion additions inside an existing trusted source can be auto-approved
|
|
|
|
This means a change like "add a missing dependency skill under an existing source" is lower risk than "replace a broad browser automation skill family", but for overlapping general skills the primary question should be retention value rather than trigger neatness alone.
|