· AI agents
DisCo and the Case for Reusable Agent Skills
Why this mattersReusable, checked procedures can help research agents build on existing work rather than repeat setup and mistakes.
An AI research agent plans, writes code, runs experiments, and explains its process in almost every language. It then implements k-fold cross-validation from scratch. Again.
Agents can resemble a brilliant theoretician who has read every textbook but has never stepped inside the lab. General knowledge travels well; the operating knowledge of a particular workflow often does not. Turning a design-of-experiments strategy into an executable plan requires knowing the data format, the machine limits, the measurement protocol, and the checks that catch a bad result.
Jianlyu Chen and colleagues’ DisCo agent addresses that gap by distilling know-how from repositories and papers into reusable skills. Its Creator constructs the skills; its Researcher uses them to solve subsequent tasks. Different operating instructions let the same underlying model take on different research roles [1, 2].
What is a skill, really?
A skill packages its operating instructions, references, and callable scripts in a folder.
calendering-thermal-conductivity/
SKILL.md
references/
derivation.pdf
parameter-tables.csv
scripts/
compute_keff.py
fit_parameters.py
plot_ushape.py
SKILL.md is the agent-facing manual: what the skill does, when it applies, which inputs it expects, how to run it, and what usually goes wrong. The references hold deeper evidence. The scripts expose procedures that the agent can call instead of reimplementing. The names above illustrate a package structure; the wrappers should point to the project's actual implementation.
Here, is the set of skills and contains routing or dependency links. An entry skill points to the relevant component: data preparation, fitting, evaluation, diagnosis, or repair. The agent opens the material it needs at each step. This progressive disclosure makes a large library usable without loading every manual into the context window [1].
Scope, ground, construct, verify
A repository, paper, or concrete problem supplies the anchor . Capabilities are grounded in evidence , packaged into a candidate skill graph, and verified with a construction record .
- Scope. Identify the capabilities worth exposing. Starting from a repository might yield “fit a thermal closure” or “run inference.” Starting from a task such as “increase energy density while keeping adhesion above a threshold” instead reveals capability gaps.
- Ground. Gather source code, documentation, examples, configurations, data definitions, and tests. A task-driven pass may need additional papers or operating procedures to fill gaps.
- Construct. Turn that evidence into instructions and callable tools. Package the routines, define their inputs and outputs, and connect the resulting skills.
- Verify. Exercise the procedures with assertion-backed cases, small fixtures, and native examples or tests. Repair failures locally and retain unresolved gaps in the construction record .
A file that describes an experiment is only part of the job. Verification has to establish that the procedure runs and that its checks detect meaningful failures. For manufacturing, useful checks include physical bounds, conservation, approved parameter windows, and a trace from each reported output to the data and configuration that produced it [2].
Creator and Researcher share the same machinery
DisCo separates skill construction from skill use while retaining the model and execution harness.
is the language model, is the harness for planning, tools, memory, and verification, and is the operating knowledge available through skills. Creator builds and verifies that knowledge; Researcher retrieves the subset relevant to its task. Repository distillation is paid for once and reused, although source changes still require maintenance.
This is useful when an agent repeatedly works on related problems. The workflow for a hyperparameter search can preserve the correct split, scoring convention, and recovery path across many runs. The next session can start from that procedure instead of reconstructing it from scattered notebooks.
DisCo on my calendering model
For this article’s 30 September update, I ran DisCo 0.2.1 in Creator mode on my calendering project, using Kimi K2.6 through OpenRouter. The anchor was the actual source repository: the electrode thermal model, data-loading routines, measured calendering states, documentation, and tests. I scoped the extraction to forward through-plane conductivity and the contact/reorientation workflow [3].
Creator identified the conductivity and inverse-porosity APIs, the family-specific calibration routines, and the distinction between gas-filled and electrolyte-filled pores. It exercised ten groups of assertions against the repository’s Python implementation on CPU.
| Check | Result |
|---|---|
| Wet coating at porosity 0.30 | Graphite: 2.559; NMC: 0.970 W/(m·K) |
| Forward/inverse porosity round trip | Recovered 0.3200000000 from a case generated at 0.32 |
| Adding a contact bridge | Conductivity increased from 0.3905 to 1.3936 W/(m·K) |
| Graphite orientation limits | 6 W/(m·K) at full alignment; 102 W/(m·K) for the isotropic orientation average |
| Repository calibration summary | Average MAPE: 31.1% for the baseline; 4.5% after calibration |
All ten check groups passed, including the Knudsen correction, isotropic NMC limit, material-family registry, and a single-family calibration. A second pass checked seven integrated results, including the contact and reorientation fits and inverse-porosity uncertainty. The generated skills and execution record keep the inputs and results inspectable.
These are small, fixed-parameter API cases. The error figures reproduce the repository’s own calibration summary and measure in-sample fit; held-out states and transfer to another chemistry still need their own evaluation. Conductivity alone can also leave contact evolution and graphite reorientation difficult to distinguish.
A small Researcher demo
I then opened a separate Researcher session with the generated skill and asked it to compare dry graphite at three porosities. It produced a short Python wrapper around the existing APIs. The sweep used an 18 µm particle diameter, solid conductivity of 25 W/(m·K), and bridge conductivity of 130 W/(m·K).
| Porosity | No bridge | Contact fraction 0.01 | Ratio |
|---|---|---|---|
| 0.25 | 0.754 W/(m·K) | 1.872 W/(m·K) | 2.48× |
| 0.35 | 0.481 W/(m·K) | 1.525 W/(m·K) | 3.17× |
| 0.45 | 0.318 W/(m·K) | 1.279 W/(m·K) | 4.02× |
Separate wet-graphite forward/inverse checks recovered all three porosities with errors below . This is a small illustrative sweep, with the script and numerical output retained alongside the skill. The fixed contact fraction demonstrates the model’s response; a recipe-specific prediction requires calibration.
That exercise made the Creator/Researcher split tangible: the first session assembled the API conventions and checks; the second used them for a new request. Reviewing the generated record still mattered. I corrected its date, check-count labels, and an unsupported total-runtime claim against the command output. Execution evidence belongs next to the instructions.
The continual-learning work supplies another recurring workflow: specify the task sequence, freeze or adapt the representation explicitly, compare IGFA with matched baselines, and record accuracy together with forgetting. A low forgetting score has little value if projection has also stopped the new task from learning. The feature regime, random seeds, retained update norm, and available capacity belong in the operating instructions.
The practical starting point is one repository and one short, executable chain. Keep units, parameter conventions, and physical checks in the skill; route to deeper references when needed. For larger Researcher tasks—comparing acquisition rules, predicting held-out states, or running an IGFA sequence—the data split, task order, and resource budget should be fixed before comparing results. The CPU checks establish that the packaged procedure runs; measuring how much time or accuracy the skill adds requires a separate comparison.
Why operational knowledge is worth the effort
The AREX-Skill authors report gains while keeping the backbone, harness, and execution budget fixed.
| Benchmark | Metric | Without skills | With skills |
|---|---|---|---|
| MLE-bench | Any-medal rate | 31.11% | 72.89% |
| PaperBench | Replication score | 29.45 | 39.59 |
| FrontierCS | Score | 70.63 | 77.14 |
| PassNet | AS score | 1.343 | 1.531 |
Those correspond to relative gains of 134.3%, 34.4%, 9.2%, and 14.0%. They are the authors’ benchmark results, rather than a measurement on my projects. The implication is nevertheless interesting: better operating context can improve an existing agent without changing its backbone [1].
For a lab, the practical starting point is a narrow domain and a few recurring tasks. Pick the repositories and papers those tasks depend on; list the capabilities actually used; collect the examples and checks; wrap the stable routines; and connect the resulting skills through an index. Then record which skills get reused, where they fail, and what still requires improvised reasoning.
Task-oriented distillation extends the same approach. When a new process objective exposes a missing capability, the agent can gather the relevant evidence and construct a skill for it. Over time, the library can hold procedures for designing experiments, analyzing measurements, and responding to drift. This gives a concrete form to the continual-learning process controller I have been sketching.
Many labs have accumulated decades of knowledge about running a process. Research agents still rediscover the same data conventions and experimental checks each time. DisCo offers a way to make those procedures explicit, callable, and maintainable. The next agent should be able to use the wheel that already exists.
Sources
- Chen, J., et al. Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills, 2026. See also the authors’ AREX-Skill library and benchmark table.
- VectorSpaceLab. DisCo Creator and Researcher workflows.
- Störk, J. Zehner electrode thermal model; calendering note and manuscript.
- Störk, J. The Geometry of Forgetting in Continual Learning, with draft manuscript.
Use your identity on another device
Save an encrypted identity file and restore it in another browser. Keep the file and its passphrase private.