What an agentic software factory is
Agents versus assistants, orchestration, isolation and guardrails, where agentic stops and autonomous starts, and what to ask a vendor selling one.
Seif Sgayer, Founder
Updated 8 min read
In one line
The word agentic says who does the work. It does not say who approves it.
On this page
An agentic software factory is a delivery loop in which AI coding agents, not assistants, do the work between a ticket and a pull request. An agent takes a task, plans, edits the codebase, runs the tests, reads the failures, tries again and reports back, inside a harness the team controls. The word agentic says who does the work. It does not say who approves it, which is a separate decision this guide gets to.
This is written for engineering leads who are being shown products and pitches with these words on them and want to know what is real, what is a label, and what to ask.
Agents versus assistants
The distinction is about who holds the loop. With an assistant, a developer holds it: they ask, read the answer, apply it, run the tests, and ask again. With an agent, the software holds it: given a goal and tools, it decides the next step, takes it, observes the result, and continues until the goal is met or a limit is hit. The same underlying model can be either. The difference is the harness around it.
| Assistant | Agent | |
|---|---|---|
| Who holds the loop | The developer | The software, within limits |
| Starts from | A prompt typed by a person | A ticket, a label or a schedule |
| Runs where | The developer's editor or terminal | An isolated runner or container |
| Can run the tests | If the developer does | Yes, as part of the loop |
| Output | Suggested code | A branch and a pull request |
| Repeatable across the team | Depends on the person | Yes, the harness is shared |
Most teams already have assistants. Many have had a developer who can use them very well. A factory is what you build when you want that developer's results from the whole team, from the ticket onward, without that developer standing in the middle.
Orchestration: what starts the work and what finishes it
Orchestration is the unglamorous part and the part most pitches skip. Something has to notice a ticket is ready, start a run with the right context, wait for it, handle failure, open the pull request and tell the ticket what happened. In a well-built factory that something is a small amount of workflow: a GitHub Actions job triggered by a label or an issue event, a webhook from Jira or ClickUp, a scheduled run for maintenance work.
The orchestration layer also enforces concurrency and order. One run per ticket. No two agents editing the same files at once. Retries capped. A queue when the team labels twenty tickets on a Friday. None of this is intelligent, and all of it is what separates a factory from a demo.
A useful test when evaluating a product: ask what happens when the agent fails halfway. If the answer is a stack trace in a log that nobody is watching, the orchestration is not finished. If the answer is a comment on the ticket explaining what went wrong and what the agent recommends, it is.
Isolation: where the agent runs
An agent with tools can run shell commands, install packages, read files and make network calls. That is the point of it, and it is also why it must never run on a developer's laptop with their credentials. Every factory run should happen in a fresh environment: a CI runner, a container, or a sandboxed virtual machine, with a clean checkout of the repository and nothing else.
- Fresh per run. The environment is created for the task and destroyed afterwards, so one run cannot contaminate the next and a mistake has nowhere to persist.
- Minimum credentials. The run holds a token scoped to push a branch and open a pull request on one repository, and the model provider key. Not a deploy key, not a production database password, not a developer's personal token.
- Secrets out of the prompt. Secrets live in the secret store and are injected into the environment. They do not appear in the task text, the rules file or the logs.
- Inspectable. The full run log, the commands the agent ran and the diff it produced are kept with the pull request, so a reviewer can see how the change came about, not just what it is.
Guardrails: the limits the harness enforces
Guardrails are the rules the agent cannot talk its way around, because they are enforced outside the model. A good factory has them at four layers.
- Repository rules. A rules file in the repository states conventions, the test command, forbidden paths and the definition of done. The agent reads it every run. This is guidance; the model can still get it wrong, which is why the other layers exist.
- Run limits. A timeout, a cap on tool calls or turns, and a budget. A confused agent stops spending at a known bound.
- Permission boundaries. The token the run holds cannot merge, cannot deploy and cannot touch other repositories. Code ownership rules route changes to sensitive paths to named reviewers.
- Branch protection. The main branch requires passing checks and an approving review. Whatever the agent produces is a proposal until a person, or a deliberately chosen policy, accepts it.
Notice that only the first layer involves the model's behaviour. The others would hold even if the agent were adversarial. That is the standard to aim for: the factory is safe because of what the agent cannot do, not because of what it has been asked not to do.
Where agentic stops and autonomous starts
Agentic and autonomous are used interchangeably in marketing and mean different things in practice. Agentic describes how the work is done: an agent holds the loop from ticket to pull request. Autonomous describes who decides it ships: nobody. An autonomous software factory is one where the agent's output merges and deploys without a human decision, which is the dark factory described in our guide to light and dark software factories.
Every factory worth the name is agentic. Very few should be autonomous, and none should be autonomous for all of a codebase at once. The sensible path is a dial: every change approved to begin with, lighter review for routine changes once the suite has shown it catches what matters, and auto-merge only for narrow, named classes of low-risk change. A vendor who presents autonomy as the default rather than the destination is describing a demo.
A practical way to hold the line: write down, per path or per label, who or what approves a merge. If the answer for any path is "the checks", that path is autonomous and should have earned it. Everything else is agentic and reviewed.
What to ask a vendor
Products in this space range from genuinely useful orchestration to a chat window with the word agent on it. These questions separate them in a single call.
- Where does the agent run, and with which credentials? You want a fresh, isolated environment per run and a token scoped to one repository. Anything that runs on a developer machine or holds a broad organisation token is not a factory.
- What starts a run? It should be an event in a tool we already use: a label, a column move, a webhook. If a person has to type into the product to begin work, it is an assistant.
- Can it merge? The right answer is no, not without our branch protection saying so. Ask to see how the product behaves when a required review is missing.
- What happens when a run fails? Look for a report on the ticket and a clean stop, not a silent retry loop or a half-finished branch.
- Where do our rules live? In our repository, in a file we can edit, version and review. Rules held only inside the vendor's product leave with the vendor.
- What do we own at the end? If the answer is a subscription, the factory is theirs. If it is workflow files, a rules file and documentation in our repositories, it is ours.
- What does it cost per pull request? Model usage, compute and the product fee, separated. A vendor who cannot attribute cost per merged pull request cannot help you decide which work belongs in the factory.
- Which of our tools does it actually connect to today? Ask for tested versus on request, named tool by tool. Roadmap slides are not connectors.
The answers also apply to a factory you build yourself, or have built for you. The Nx team's argument that a software factory is a workflow, not a product is the right frame: whatever you buy or build, it should sit inside the tools you already run, and you should be able to read every line of the glue.
That is how we approach our software factory setup engagements: the workflow, the rules file and the gates are built in the client's own GitHub and ticket tool, and the team can change any part without us.
Common questions
What is an agentic software factory?
A delivery loop in which AI coding agents do the work between a ticket and a pull request: plan, edit, run tests, retry and report, inside a harness the team controls. Agentic describes who does the work. Whether a person approves the merge is a separate choice, and most teams should keep it.
Is an agentic software factory the same as an autonomous one?
No. Agentic means an agent holds the loop from ticket to pull request. Autonomous means the output merges and ships without a human decision. Every useful factory is agentic. Autonomy is a per-path setting earned by a trustworthy test suite, not a default, and it is what others call a dark factory.
Do we need a platform to run AI agents on our backlog?
No. A GitHub Actions workflow, a rules file in the repository, branch protection and the agent of your choice are enough for one repository. Platforms can add orchestration, dashboards and connectors, which is worth paying for at scale. Start with the workflow so you know what the platform would be replacing.
Which AI agent should a software factory use?
Whichever your team can run in an isolated environment with scoped credentials and a readable log, and whose provider's data terms you have read. Claude Code is one we have tested end to end inside GitHub Actions. The harness matters more than the model: a good loop around a decent agent beats a great agent on a laptop.
More guides
The other four notes in this series.
Get your software factory set up
We start with one repository and one workflow. Once it works, we add more.