Every organization operates at two speeds. There is a slow-thinking layer where you figure out what to build: strategy, experimentation, hypothesis formation, judgment calls about which approaches to pursue and which to abandon. And there is a fast-thinking layer where you execute: well-defined work, clear success criteria, deterministic pipelines.
Humans have always done both. The interesting question is what happens when agents can do the fast-thinking layer autonomously. The answer, I think, is that agent orchestration splits along that same divide, and the two sides need different solutions. Most of the conversation right now conflates them, which is why the space feels so confused.
Tasks and Goals
A task has a known shape. Ticket in, PR out, review, merge. The success criteria are defined before the work starts, and the pipeline runs one direction. An agent picks up a Linear ticket, opens a PR, gets reviewed, iterates, and ships. That is a job.
A goal does not decompose cleanly. “Increase activation rate by 20%” has no ticket shape. It needs research, hypothesis formation, experiment design, cross-functional coordination, and calls about which ideas to chase. The orchestration I have working is all task-shaped.
The Dark Factory
Manufacturing has a term for a fully automated plant that runs without human workers: a dark factory. Lights off, machines running, parts moving through the line around the clock.
The software version is an agent taking a ticket and producing a PR, an architect agent reviewing the approach, a QA agent testing the branch, a code reviewer gating the merge. Each step is a discrete compute unit with defined inputs and outputs. That is what my /implement pipeline is, and what every Claude PR check and every Codex agent picking up a ticket already does at a single-step scale.
The economics work even when the system is wildly inefficient. Gas Town reportedly burns something like ten times the tokens a single focused agent would spend on the same work. That does not matter if the output is still far cheaper than an engineer, because then efficiency is an operations problem. Tokens keep getting cheaper and intellect per dollar keeps going up. Even if the concept is in its early stages, if you can get high-level direction executed with some quality and overarching determinism, the economics of replacing teams with agent pools are compelling. That is where the one-person, multi-million-dollar company talk comes from.
Stateless Per Task
Inside the factory there are two competing shapes for a worker.
The persistent session runs one agent all day, with Slack messages, ticket updates, email, and PR comments all funneling into a single long-lived context. It has natural continuity. It also spends tokens on context irrelevant to the current task, and the context window is finite, so you compact aggressively and end up doing memory management whether you planned to or not.
The stateless-per-task model spawns a fresh agent for every incoming signal and loads identity, learnings, and memory as context at invocation. It is more efficient per task, and it scales horizontally because you can run any number of sessions at once.
The catch is memory. If every session is fresh, continuity between tasks depends entirely on how well you have externalized state. Everything an agent knows outside its current task has to live in memory rather than in the conversation. I am probably underestimating how hard that is. But with MCP, new memory features, and the cottage industry of memory supplements, I think it gets solved well enough, and I think that outweighs the inability of persistent sessions to scale.
The container world already landed here. Kubernetes favors interchangeable workers and pushes the state that has to survive into purpose-built stores. The state that matters here is learnings, memory, PR feedback patterns, and codebase context. The agent is the compute, and the orchestration layer holds everything durable.
Where the Hard Part Lives
The dark factory is the tractable subproblem. The hard problem is goal decomposition: turning a goal into a set of tasks a dark factory can execute. It cannot be another station on the factory floor: every station needs its success criteria defined before the work starts, and a goal is the thing that arrives without them.
This is the difference between algorithmic and heuristic work. Algorithmic work wants tight processes, clear handoffs, and minimal coordination. Heuristic work wants something closer to Steve Yegge’s VP metaphor: someone who reviews nobody’s output but coordinates across efforts, resolves conflicts, and decides when to pivot.
The quality of the decomposition determines whether the factory works at all. Clean, orthogonal tasks with clear success criteria, and the factory just runs. My /implement pipeline gets this in engineering because the planning phase assigns files to agents with no overlap. Good decomposition at the top eliminates coordination at the bottom. Messy decomposition, with tasks that interact and share state and sequence against each other, puts you back in coordination hell.
Above the decomposition sits the part that is still human as far as I can tell: which goals to pursue, which tradeoffs to accept, when to change direction.
The interface between goal decomposition and the factory floor is where most organizational dysfunction lives, in human orgs and agent orgs alike. In human orgs it shows up as sprint planning and roadmaps. In agent orgs it is the quality of the goal-to-task handoff. The dark factory will end up commodity infrastructure. The decomposition sitting on top of it is not infrastructure, it is the product.