An AI lab

Private intelligence, built where no one is looking.

Most agent research assumes the work can leave the building. Ours assumes it can't. That meant rebuilding the harness, the compute, and the permissions underneath it.

A
Bridge
Control plane
Shipwright · above the line
B
Harbor
Main deck · 0 fathoms
C
Engine Room
Below decks · 30 fathoms
Long-horizon agentsHarness technologyLocal computeAgentic security
Why the yard exists

The work is leaving the desk.

An agent that runs for six hours is not a chatbot on a longer leash. It opens files no one listed, spends money no one approved line by line, and reaches hosts that were never in the plan — most of it after the person who asked has gone home.

The usual answer is to run it somewhere else: someone else's data center, someone else's jurisdiction, someone else's audit log. For most teams, that is the right answer.

The organizations with the most to gain from a six-hour agent are the ones that can't use one. Not won't. Can't.

So we build the other way.

Four things that must be accounted for
Every file opened
Every dollar spent
Every host reached
Every decision made unattended
The ship

One hull, three decks.

Three products, stacked the way a ship is stacked. You stand on the top deck. The other two are why it floats.

IHarborMain deck · 0 fathoms

Where you speak to it.

A quiet room. Ask, hand off, come back to finished work. It picks up where you left off on any device, and nothing leaves your dock.

Fig. I — the two forms
AUDone line in the conversation
a full berth for the thing itself

A hundred results don't become a hundred walls of text. Each one gets a line where it happened, and a place you can go back to.

Waterline

Everything below this line runs on hardware you can walk up to.

IIBridgeControl plane · above the line

The yard where it is built.

Every tool, every hand, every permission, accounted for. Work moves through it in the open, and what it touches stays yours.

Fig. II — two phases, one wall
Setup
reaches out, by allowlist, once, with your consent
GATE
Agent
works alone, egress closed, every reach recorded

The allowlist is written while you are watching. The model never edits it, and has no way to ask.

IIIEngine RoomBelow decks · 30 fathoms

Below decks, where it draws its power.

The part you never see. Enough compute to matter, close enough to trust. Nothing leaves the ship unless you send it.

Fig. III — the ledger
Ran on this machineall of it
Held on this machineall of it
Hosts reachednone
Left the networknothing

A measurement, not a policy — which is why it can also tell you when something did leave.

Research

Four problems we are working on.

The dashed line under each is the part we haven't solved.

01

Long-horizon agents

Context windows got bigger. Memory did not. Four hours in, an agent is mostly working from a summary of itself, and the summary is where the errors live.

Open problem

What should an agent refuse to compress?

02

Harness technology

Benchmarks measure the model. Deployments fail on everything around it: tool schemas, retries, timeouts, the interrupt that lands mid-write. Most of the distance between a good demo and a working system is here, and almost none of it is in the weights.

Open problem

Which failures belong to the model and which belong to the harness? There is still no clean way to tell them apart.

03

Local compute management

Running a serious model on hardware you already own is a scheduling problem before it is a hardware problem. Weights that don't fit, batches that never fill, eight people and one GPU.

Open problem

What is the real floor — how much iron does a building actually need before this stops being a compromise?

04

Agentic security

Prompt injection is the famous one. The quieter problem is an agent that had permission for every individual step, and should not have had it for the sequence.

Open problem

How do you prove what an agent did not do?

Instruments

What we measure.

We would rather show four empty boxes than four numbers we can't reproduce. They fill in when the runs do.

Endurance
hours
Unattended running before a human decision is required
Completion
per cent
Long tasks finished without intervention
Egress
bytes
What left the network over a full run
Throughput
tok/s
Sustained on hardware the customer already owns
Evidence

A run, start to finish.

Run 0417, plotted from its own event log: six hours on sealed hardware, compressed to forty-five seconds, including the parts where it sits and waits.

Drawn from the event log, not captured off a screen. Each hairline is one action; teal marks a decision rather than a read. The gaps are real, and none of them were cut.

Berths are few, and we open them by hand.

If you have a constraint we should hear about

We take on a few partners at a time, and we choose by whether the constraint is real. Tell us yours.

Request a berth
Or stand off and watch

Notes from the yard when there is something worth reading. No cadence, no newsletter.