⚒ FORGE v1.0.0

your harness, your rules
Server
engine (27B)
VRAM
GPU util
Sampling
Thinking
enable_thinking
preserve_thinking
Qwen warns: low effort in agent loops can cost MORE time overall via failed attempts and retries.
Agent
auto-approve (YOLO — deny rules still apply)
System prompt
Replaces the entire system prompt. ~90 tokens by default — nothing hidden.
Permissions
saved ✓
Agent
Arena
Models
MCP
model fleet 0/0
FORGE — give the agent a task below
Enter send · Shift+Enter newline · Esc abort
settings apply per message — tune freely between turns
turn speed tok/s last gen tok total out tok context in tok tool calls
Fleet
Lanes run concurrently via vLLM continuous batching. A profile with an endpoint set can point at a second server (e.g. llama.cpp on :8001).
Select profiles, enter a prompt, hit Run.