ABVX Lab logo ABVX Lab
learning layer

ABVX explainer / agent learning

What actually learns in an AI agent?

Most practical agent improvement does not mean changing model weights. It means routing lessons into the cheapest auditable layer that will improve the next run.

01

Model layer

Weights, fine-tunes, adapters, and provider behavior. Powerful, but expensive, slow to audit, and rarely the right first move for workflow learning.

  • Use for stable capability gaps.
  • Avoid for one-off preferences.
  • Needs stronger evaluation.
02

Context, memory, skills

The ABVX default: repo facts, operator conventions, checklists, SKILL.md workflows, and progressive-disclosure behavior.

  • Cheap to read and revise.
  • Reviewable in pull requests.
  • Best for repeated agent work.
03

Harness and workflow

Scripts, gates, CI checks, fixtures, judge rubrics, and golden outputs. Use this when prose keeps failing or regressions must be caught.

  • Turns judgment into checks where possible.
  • Bounds loops with budgets and stop rules.
  • Prevents repeated failures from returning.

Routing rule

Promote only when the next layer earns its cost.

A useful lesson should not automatically become a global prompt or skill. Sometimes it is just a memory note. Sometimes it should become a script. Sometimes it should become an eval.

Prompt/session: one-off correction, no durable value.
Memory or durable docs: stable facts, preferences, repo setup, verification paths.
SKILL.md or checklist: reusable behavior with trigger discipline and anti-patterns.
Script/tool or eval: deterministic repeated work or failures that must not regress.

Copyable decision record

Use this after an agent run teaches you something.

Learning candidate: <one sentence>
Evidence: <trace, PR, review, failure, success pattern>
Chosen layer: prompt | memory | durable-doc | checklist | skill | script | eval | reject
Why this layer: <short rationale>
Artifact to update: <path or destination>
Verification: <how we know the learning helps>
Rejected higher layers: <why not skill/script/eval/etc.>
Next action: <one concrete edit or no-op>