Why we built our coding agents on an office desktop
Here’s why we built coding agents on an office desktop, how it works, and how it helps us ship faster.
Manu Mahajan
Head of Engineering

At Seapoint, engineers can assign a Linear ticket to a service called Seapoint Coding Agents (SCA). Shortly after, a pull request is ready for review, replete with testing evidence: screenshots, videos of browser sessions, URLs with a live preview of the code.
The whole system runs on a desktop machine in our office, and engineers can assign work to hosted agents from their local development environment or via Linear.
Since March 2026 the pipeline has merged over 400 PRs, and last quarter autonomous agents wrote the fix for about 60% of resolved customer bugs. As a result, median time for an issue to go from reported to resolved was 2.4 days, and some features shipped in as little as 24 hours after being reported.
Software factories like these are where we think the industry is heading, and by experimenting in this way we’re keeping ahead of the curve. Here’s why we built coding agents on an office desktop, how it works, and how it helps us ship faster.
Building our agents: how we started
In the early days of using Claude Code, Seapoint founder Sean and I wondered if agents could run while our laptop was closed for prolonged periods, or even overnight.
We started by running multiple tmux sessions on laptops and remote developer machines on fly.io, before buying a small desktop machine to keep in the office for developers to run remote Claude sessions. RAM was expensive, but one machine between five developers made the expense worth the while.
The remote sessions worked for ad-hoc tasks, but we soon realised we could automate small coding tasks of a similar nature. So, we built a simple API on top of Claude’s (then new) agent SDK, with a scrappy UI to interact with it.
Getting started: building the foundations
Before we could deploy agent fixes we first needed to build the right infrastructure, such as:
- Shared context. All coding guides, prompts and skills shared by the team are situated in our repository under
.agents/promptsand are used by agents to plan and build tasks. - Sandbox VMs. Every time an agent picks up a task, it creates a new git workspace, pulls in code and developer tools, and runs the code in a sandboxed docker container under a unique URL only developers could interact with. When building this, we also added egress rules, mock services and dummy data.
- Manual overrides. More often than not, an agent does 80-90% of the work, but the PR would need a developer to make changes. To do this, we made it easy for engineers to connect VsCode or Claude CLI to the remote VM via SSH and work directly from there, without pulling code to your development machine.
- Agentic testing loops. We built a verification system that could open a browser and test the code like a user, running operations and posting results as screenshots and videos to generate PRs.
This isn’t a model or agent loop, but rather the infrastructure that supports it: a task queue, container supervision, a live dashboard and a GitHub App that opens pull requests.
Use cases
Fixing tests
As a first use-case, we put it to work every time our end-to-end tests failed. These tests used Playwright to log in to our staging website with a user and carry out operations to make sure our core functionality worked.
The problem was, these tests would break a lot. We were testing a product in constant development, which meant that any time copy or UX was updated, the test would fail.
Because this type of testing infrastructure has always been difficult to maintain, it’s been a great use case for agents. For the most part, if you can tell a page has been updated because of a UX change and not because of a bug, so can an agent.
So whenever a test would fail, a Github Action would call this API endpoint to fix the tests. If it couldn’t, it would create a PR and ping the oncall on Slack to review it, and since the code wouldn’t be customer-facing, we could afford to be more relaxed with code reviews.
Solving customer bugs
As Engineering Manager, I’m always trying to find bottlenecks. When you’re building fast with a small team, a bottleneck can divert attention and slow shipping rates down.
Before we introduced our office coding agents, one of our biggest bottlenecks were bugs. Complaints would come in from both customers and stakeholders that bugs and small feature requests weren’t being fixed fast enough, so engineers would be pulled away from their main projects. This context switching shifted focus and slowed shipping, making small fixes costly.
In my previous role at Stripe, these small requests would be handled by a runner, an engineer who would be responsible for fixing bugs, answering questions and managing deploys to production. At Seapoint, our in-house coding agents would become our runner.
Taking inspiration from OpenAI’s Symphony, we built our own Linear-based workflow on our desktop:
- The Product team triages the tickets and marks them as
ToDo - SCA assigns tickets in
ToDoto an agent that produces a plan - An engineer reviews an approves the plan
- A PR is issued on Slack
To keep the pipeline moving, an oncall engineer reviews plans, code and merges fixes quickly.
Agent toolkit for Developers
Developers can choose which workflow they use for development, and as our pipeline has evolved we've added a new tools they can use on SCA.
- /prototype skill launches a sandbox VM with a new workspace on the server. Developers, designers and PMs can quickly build MVPs or features to share with stakeholders without worrying about shippable code. This allows us to iterate quickly and work with prototypes instead of static Figma designs.
- /delegate-to-sca skill enables developers in a session with Claude or Codex to delegate a task to a coding agent running on the server. When the code is ready, the developer can preview the deployed code in a sandbox, connect to the remote machine from their IDE or CLI, and work on further improvements if needed.
As SCA continues to evolve as a toolkit for our engineers, we have plans to build more. Here's how we think of the main system components today:
Why we built instead of bought
At first, we tried to buy. The market is crowded with strong contenders (such as Factory and Cognition, valued at $5B and $48B respectively) making the same promise: ticket in, pull request out. But for every one we tried, three requirements were never met.
First, we ship features as stacks of small, dependent pull requests using a tool called Graphite. Every vendor we tried produces one large pull request, making it harder for engineers to review code.
Secondly, we handle sensitive data. Customer banking data only leaves our network if called by the model, and most external vendors host code on the cloud. By outsourcing this data to an external vendor, we were adding an additional security risk.
Finally, our engineering conventions and files exist in our own repository, with both team members and agents referring to them on a regular basis. Since all vendors have their own repository, we’d no longer be able to access the files, becoming wholly reliant on the output without knowing how we got there.
As a result, these products are a poor fit for a team that cares about security and wants to avoid unnecessary complexity.
Why our agents run on an office desktop
As a startup, running our agents on an office desktop makes the most sense.
For one, a modern server grade machine can handle the volume of what we run for a fraction of the cost. We have up to 50 containers running at any given time, and beyond electricity and model tokens we pay practically nothing per task.
It’s also better to host internally from a security perspective. We’re dealing with a lot of confidential information, and on our own machine both code and data stay inside our network. The server sits behind Tailscale, every container uses an encrypted volume, and we decide which external services a container may talk to. If an agent goes rogue it cannot widen its own network access, and every blocked attempt to escape is logged, keeping sensitive data secure.
What we’d tell a team facing the same call
Invest in the infrastructure that will accelerate development with agents. This can include creating shared context, building a place to run code safely, and storing evidence a reviewer can trust.
Start with a job where mistakes are cheap and put your conventions in the repository, so people and agents follow the same rules. Make every piece of agent work show its evidence, make sure there’s a person to make the judgement calls, such as approving a plan or merging a change. Measure it, and revisit the decision every few months.
For now, one desktop in the office has taken us a long way, and we suspect it will take us further still. If this is the kind of engineering you want to do, we’re hiring.
