What It Takes to Let an Agent Finish the Job

Someone mentioned a problem in a recorded meeting. An agent picked up on it, pulled a few frames from the recording to capture what the person had seen, and investigated. It found the bug, made the change, and verified the fix.
Nobody had to explain the problem again. Nobody had to walk the agent through the next step.
That’s a useful way to describe what we’ve been building at Gearflow. We want someone to report a problem in the course of their work and have that problem get resolved, without an engineer having to carry it through every stage.
Getting there means answering a lot of practical questions. Where does the agent find the history behind a decision? Can it inspect the data involved? How does it know whether it’s allowed to merge? What happens when its session ends halfway through?
Our answers live in what we call the agent harness: the shared knowledge, tools, permissions, and operating rules around the model. Here’s how those pieces fit together.
Context: the history behind the work
An agent needs to know what it’s walking into.
A bug report rarely contains everything needed to fix the bug. Some of the explanation lives in a previous issue. Some lives in a conversation with a customer. Some lives in an architectural decision somebody made months ago for a reason that still matters.
An engineer knows to go looking for those things. We give agents the same expectation, along with access to the record.
Every agent starts with a shared, version-controlled workspace containing our architecture notes, runbooks, product decisions, operating instructions, and lessons from previous failures. For the task itself, it reads the issue history, related work, relevant chat threads, meeting notes, and customer context.
That changes the quality of the work before anyone writes code. An agent can discover that a proposed solution conflicts with an earlier decision, or that a seemingly small change touches a workflow another customer depends on.
It also gives the agent something meaningful to check at the end. Passing tests is part of finishing. So is making sure the change addresses what the person originally meant.
We even keep writing expectations in that shared context. Lead with the conclusion. Use short sentences. Cut the filler. We read agent output throughout the day, and every extra paragraph is something a person has to process before deciding whether the work makes sense.
Capabilities: access to evidence
The next requirement is access to evidence.
If an agent has to ask an engineer to run every query, retrieve every log, and take every screenshot, that engineer is still doing much of the investigation.
Our agents can query production databases and the warehouse through read-only connections. They can inspect traces, logs, and errors. They can run the application with anonymized data shaped like production data, use the interface, and capture screenshots.
That lets them follow an investigation through. How many records are affected? Which code path produced the error? Can the problem be reproduced? Does the proposed fix work with the kinds of data we actually have?
The same access supports verification. A data answer comes with query results. A UI change comes with screenshots. The agent has to bring back evidence that someone else can inspect.
Trust: what an agent is allowed to do
Of course, being able to investigate a problem doesn’t automatically mean being allowed to change everything involved.
We use explicit trust levels, recorded as labels on issues. Those levels range from read-only investigation through permission to merge. Agent merge is our default for issue work, and the person filing or approving the task sets the level. The agent records whose authorization it acted on.
The entry point matters, too. A chat notification can trigger a read-only investigation and a reply. A labeled issue can authorize an agent to claim a working environment and deliver a fix. Meeting notes provide another way for work to enter the system.
Underneath, these follow the same pattern: an event starts a session, the session loads the shared context, and the agent works within the permissions for that task. Adding another entry point is easier because those operating rules already exist.
We’ve also had to be specific about when agents should ask questions.
Our rule is to search the record first. If the answer is already in an issue or a previous discussion, the agent should find it. When a question is necessary, it should come with a recommended default. For a reversible action within its authority, the agent states when it plans to proceed with that default. For an irreversible action requiring an answer, it waits.
That keeps human attention available for decisions that actually need it.
Many agents at once
Once several agents are working at the same time, the problems become familiar software infrastructure problems.
They need separate working copies. Otherwise, one agent’s changes can break another agent’s build. Each gets an isolated checkout, database, and application server. A small lease system tracks which environments are occupied, and agents release them when they finish.
They also need a reliable way to hand off work.
A task can outlast any individual session. Context windows fill up. Sessions end. A pull request might wait days for a reply. We treat that as a normal part of the workflow.
The issue tracker holds the durable task record. Before starting, an agent reads the description and comments. Before stopping, it records what it completed, what remains, and where to find the relevant work. Another session can pick up from that record without depending on the previous session’s local memory.
Waiting works similarly. An agent waiting for CI or a human response schedules a wake event and ends its active session. When the event arrives, the harness resumes the conversation. There’s no need to keep an agent process running just to watch for a build result.
A harness that learns from its work
The harness also needs a way to learn from the work passing through it.
When an agent discovers a deployment constraint or a recurring data trap, it opens a pull request against the shared workspace. Once that knowledge is incorporated, future agents and new engineers can find it.
Human corrections belong there, too. A correction that stays in one conversation is easy to lose. Writing it into the operating instructions gives the next session a chance to avoid the same mistake.
Those instructions need maintenance. A rule can solve one failure and later get in the way of legitimate work. When that happens, the agent opens an issue against the harness. Agents and engineers then work through a proposed change, with an engineer approving the update. That gives us a way to repair a bad rule without making exceptions silently.
Gates that only get stricter
Verification follows the same principle: make the standard explicit, then require evidence that the work meets it.
Agent changes go through our CI suites, automated code reviews, and resolution of review comments. The agent follows the results, fixes failures, and handles feedback until it reaches the merge or handoff allowed by its trust level.
We’ve also used agents to make those checks stricter.
Adding a static analysis rule often creates a backlog of existing violations. Clearing that backlog takes time, which makes a useful check easy to postpone. Agents can do that cleanup. We add the rule, resolve the violations, and make it required. Compilation warnings now fail the build.
We keep those checks in place when they become inconvenient. A failing gate is work to resolve before the change goes through.
This is what makes more autonomy practical. Each required check gives us another concrete reason to trust a change.
What this changes for us
The effect on our team is visible in ordinary work. Product and support can report a bug and get a tested pull request, frequently the same day. Routine investigations and fixes can move forward while engineers spend time on design, review, and decisions that need their judgment. Work can continue after hours, with a written record ready for the next morning.
The knowledge stays with us, too. Each useful discovery becomes something the next person or agent can start from.
If we were building this again, we’d begin with the information an engineer needs to do a good job: the history, the systems to inspect, the boundaries of their authority, and the evidence required to call the work finished.
The model remains replaceable. When we switch models, the new one inherits that working environment.
That’s the part we keep investing in, and the work itself does much of the investing. Every finished task leaves something behind: a lesson in the shared context, a rule from a person’s correction, a check that is now stricter. The next agent starts from all of it. Each agent that finishes a job leaves the harness better for the one after it, whichever model runs inside it.






