Research paper / September 2026

WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents

WEFT builds diverse interactions, improves the system through execution feedback, and trains at verifiable task boundaries.

Scale the interaction system, not just the environments.

Useful tool-use training depends on environments, executable tasks, agent harnesses, and evaluators working together. WEFT scales, improves, and learns from this entire interaction system.

The WEFT framework

Explore the three stages

Build stateful environments and verifiable tasks, then route tasks through different interaction settings and harnesses.

Explore construction
01 / CONSTRUCTBroader interactions

A richer source of training interactions.

Executable, verifiable tasks are hard to obtain. WEFT expands their training value through environment breadth, task complexity, and interaction diversity.

WEFT pipeline: stateful environment synthesis, cross-MCP task composition, multi-view and multi-harness rollouts, and verification and revision.
Construction and execution-driven self-evolution.

Environment breadth

Synthesize and execution-test tool environments with persistent database and workspace state.

Task complexity

Compose certified atomic tasks into cross-MCP workflows: a median of 20 atomic-task turns, with more than half spanning multiple domains.

Interaction diversity

Present the same task as a complete brief (Agentic) or progressive requests from a simulated user (SimUser), then route it to multiple harnesses that retain their native execution behavior.

Executable MCPs
8,172
Tools
64,755
Atomic tasks
41,695
Composed tasks
11,884
Five panels showing domain coverage, environment and task counts, per-MCP capacity, task length, and cross-MCP composition.
Environment and task statistics.
02 / EVOLVEMore reliable feedback

Let execution improve the system, too.

Failures can originate in the policy, environment, task, or verifier. WEFT uses traces and state changes to identify the cause: retain valid tasks, revise faulty components, and revalidate them with fresh rollouts.

Three self-evolution rounds improve Toolathlon-Verified by 5.25 points, AutomationBench by 4.33 points, and Claw-Eval by 3.65 points, while tool-call errors fall by 0.80 points.
Three self-evolution rounds; fixed tasks and rollout budget.
03 / TRAINStable post-training

Learn at verifiable task boundaries.

Long workflows need reusable progress, reliable verifiers, and local rewards. WEFT organizes sampling and learning around atomic-task boundaries.

Three connected training mechanisms: prefix-preserving rejection sampling, task–verifier consistency filtering, and atomic-turn reinforcement learning.
Prefix-preserving sampling, consistency filtering, and atomic-turn RL.
01 / SFT

Preserve the prefix

Retain verified prefixes and their database/workspace states; retry only the failed atomic task from its starting state.

02 / FILTER

Check task–verifier consistency

Strict verifiers may reject correct solutions; loose ones may reward failures. Before RL, retain tasks whose executable scores agree with independent rubric judgments under both Pearson and Spearman checks. Online training uses executable rewards.

03 / RL

Assign local credit

Sample candidates from identical histories and environment states. Compute advantages from the current atomic task’s binary reward, updating only that turn’s policy tokens.

The execution layer

MegaMCP: shared services, private state.

Tool services are shared outside agent sandboxes; each rollout keeps a private database and workspace. Snapshots restore identical states for retries and branching.

Complete MegaMCP architecture: registry, private sessions, logical state isolation, shared worker processes, checkpoint verification, and snapshot recovery.
Shared tool services, isolated rollout state.
EVALUATIONWhat changes in practice

From tool calls to long-horizon workflows.

We evaluate three base models across five benchmarks. WEFT-8B and WEFT-14B lead the evaluated same-size environment-scaling baselines on BFCL V4, τ²-Bench, and Claw-Eval aggregate scores.

WEFT-14B vs. Agent-World-14B

BFCL V4
+6.41pp
τ²-Bench
+2.23pp
Claw-Eval
+12.27pp
Complete main results table with all compared models, BFCL V4 categories and aggregate scores, tau-squared Bench categories and aggregate scores, and Claw-Eval scores.
Main results (%) across model sizes.

On longer workflows, WEFT-35B-A3B reaches 45.99% on Toolathlon-Verified and 26.50% on AutomationBench.

Continue reading

The full story, in the paper.

BibTeX / arXiv export
@misc{mao2026weftscalingtooluseposttraining,
      title={WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents}, 
      author={Bo Mao and Hang He and Linting Wang and Lizhi Lin and Maosen Zhou and Guanming Liu and Jinxiu Liu and Tianyu Huai and Chaoyun Zhang and Bingxuan Li and Kepeng Lei and Guanting Dong and Zhou Shao and Rui Zheng and Hang Yan and Jie Zhou and Chengcheng Wan and Tao Gui and Liang He and Xipeng Qiu},
      year={2026},
      eprint={2609.36887},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2609.36887}, 
}

Paper figure

Select “Actual size” for detail. Press Esc to close.