---
title: "How I'm Building Tourist: A Cloud Coding Agent That Lives in a City"
description: "Tourist's city foundation is running. Here is the build plan for the cloud coding agent, scoped memory, dynamic context, GitHub delivery, and learning that should follow."
date: "2026-09-11"
updatedAt: "2026-09-28"
author:
  name: "Harsh Sinha"
  url: "https://www.harshsinha.dev"
  sameAs: ["https://x.com/sinhaharsh12","https://www.linkedin.com/in/harshsinha12/","https://www.github.com/harshsinha-12"]
canonical: "https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city"
markdown: "https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city/article.md"
tags: ["AI Agents","Context Engineering","Memory","Systems Design"]
---

> AI-readable source for [the published article](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city). Interactive components are preserved as MDX, and their structured datasets are included at the end.

## Section links

- [The problem I am actually trying to solve](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#the-problem-i-am-actually-trying-to-solve)
- [What Tourist is (and is not)](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#what-tourist-is-and-is-not)
- [Two tracks, not six products](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#two-tracks-not-six-products)
- [Track 0 — city foundation and public reports](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#track-0-city-foundation-and-public-reports)
- [MVP 1 — solo agent to PR, plus trajectories (Gate A)](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#mvp-1-solo-agent-to-pr-plus-trajectories-gate-a)
- [The stack I am starting with, and why](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#the-stack-i-am-starting-with-and-why)
- [Context: copy Cursor’s discovery model, then make retrieval a learned decision](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#context-copy-cursors-discovery-model-then-make-retrieval-a-learned-decision)
- [Search is hybrid, not “embed the repo into the prompt”](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#search-is-hybrid-not-embed-the-repo-into-the-prompt)
- [Long tool output becomes a file](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#long-tool-output-becomes-a-file)
- [Tools and skills are discovered, not preloaded](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#tools-and-skills-are-discovered-not-preloaded)
- [Summaries must not destroy history](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#summaries-must-not-destroy-history)
- [The extra step: context selection is part of the policy](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#the-extra-step-context-selection-is-part-of-the-policy)
- [Memory: persistent, scoped, and hostile to pollution](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#memory-persistent-scoped-and-hostile-to-pollution)
- [Learning: around the model, not of the model](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#learning-around-the-model-not-of-the-model)
- [The city is the report, then it becomes the runtime](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#the-city-is-the-report-then-it-becomes-the-runtime)
- [How I will implement it in the repo](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#how-i-will-implement-it-in-the-repo)
- [Original first-month plan](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#original-first-month-plan)
- [What I am afraid of](https://www.harshsinha.dev/articles/building-tourist-a-cloud-coding-agent-in-a-city#what-i-am-afraid-of)

## Author profiles

- [Website](https://www.harshsinha.dev)
- [Twitter](https://x.com/sinhaharsh12)
- [LinkedIn](https://www.linkedin.com/in/harshsinha12/)
- [GitHub](https://www.github.com/harshsinha-12)

<ArticleImage
  src="/assets/articles/building-tourist-a-cloud-coding-agent-in-a-city/pixel-art-software-development-island-city.webp"
  alt="Pixel-Art Software Development Island City, an isometric island with labeled districts for code, tools, research, testing, review, data archive, and Merge Harbor."
  caption="Pixel-Art Software Development Island City"
  width="1600"
  height="900"
  preload
/>

Most coding agents still behave like a very smart intern who forgets the office every morning.

You describe a bug. The agent searches, finds `auth.ts`, patches it, runs tests, and opens a pull request. Tomorrow you describe a related bug. It searches the same folders, rereads the same README, retries the approach that already failed on Tuesday, and dumps half the repo into the prompt “just in case.”

The model is not the bottleneck anymore. The session is.

I usually write about systems I have already poked at. **Tourist** now has a working city foundation. The larger product I am building looks like this:

```text
Connect a GitHub repo
→ bring your own OpenAI key
→ describe the work
→ a cloud agent inspects the codebase, edits and tests in an isolated sandbox
→ a pull request shows up
```

The longer version is the actual product. Most agents are still chat sessions with tools. Tourist is an attempt to make something closer to a persistent engineering environment: memory across tasks, context that is discovered instead of dumped, learning from outcomes, and a visual world that makes autonomous work inspectable.

This post is the build plan and a progress marker. The first city and report slice works. The cloud agent loop, persistent memory, tool creation, and learning are still work ahead.

<Callout title="Current checkpoint — September 2026">
  The repository runs a versioned city/report protocol, a deterministic fixture-to-city generator, an interactive isometric city, file-specific building components with a generic fallback, a building inspector, and a fixture PublicReport with file, test, and PR anchors. You can run the fixture locally. It is not yet proof of a cloud agent opening a PR or of live agents moving through the city.
</Callout>

<Callout title="The bet">
  The model is not the product. The harness, memory, retrieval, sandbox, delivery loop, and visibility are. If the city is only a skin, the project failed. If the city is how a run becomes understandable, the visualization is doing engineering work.
</Callout>

## The problem I am actually trying to solve

Coding agents got good fast. They can search a repo, edit files, run tests, and open PRs. What is still weak is everything around the session:

- they rediscover the same architecture every time
- they forget the approach that already failed last Tuesday
- they preload too much context, then drown in it
- they have a fixed toolbelt
- they do not systematically get better from merged PRs, reverted PRs, flaky tests, or “please don’t do it that way”
- when they work in the cloud, you get a log stream, not a place you can look at

I have already written about two pieces of this: [memory as a coordination layer](/articles/self-improving-agents-memory-missing-layer), and [agents that know when to stop](/articles/building-ai-agents-that-know-when-to-stop). Tourist is the project that would have to survive those constraints in production, then make the work visible.

Cursor’s own writing on this is unusually clear. In *Dynamic context discovery*, they argue that as models get better as agents, you should give them **fewer details up front** and make it cheap to pull the rest. Their contrast is static context versus dynamic context discovery. The claim is not just “saves tokens.” They say it can also improve quality by keeping contradictory or irrelevant material out of the window ([Katz, 2026](https://cursor.com/blog/dynamic-context-discovery)).

That is the first architectural tenet I am copying:

> Huge available context. Tiny active context.

The rest of Tourist is what happens if you take that seriously, then add persistence, GitHub-native delivery, and a city as the reporting surface.

## What Tourist is (and is not)

The user flow for v1 is boring on purpose. Connect GitHub, paste an OpenAI key, describe a task, wait. The interesting work is under that loop.

<Mermaid
  caption="v1 is a GitHub-native loop, not a chat transcript."
  chart="flowchart LR; A[Connect GitHub repo] --> B[Bring OpenAI key]; B --> C[Describe a task]; C --> D[Cloud agent in a sandbox]; D --> E[Commit, push, open PR]; E --> F[PublicReport city snapshot]"
/>

Under that loop I am building four things that most chat-wrapped agents skip:

<DataTable
  dataset="whatTouristIs"
  caption="The PR loop is the product. These four layers are why it is not just another wrapper."
/>

What I am not building, at least not first:

- a new foundation model
- online PPO against GPT
- Kubernetes as a personality trait
- a custom sandbox runtime
- a 40-agent swarm
- a VS Code clone
- a claim that “the city is alive” before a solo agent can open a real PR

<MarginNote>
  The city is the brand. The PR loop is the product.
</MarginNote>

If I confuse those, I will spend three months making buildings prettier while the agent still cannot ship.

## Two tracks, not six products

The failure mode for a project like this is obvious. Memory, multi-agent, tool factories, RL, and a 3D world are five startups. I wrote quality gates into the PRD so I cannot “just quickly add a swarm.”

There are two tracks. The first Track 0 slice is runnable; the solo PR loop remains the next major gate. Everything else is gated.

<Mermaid
  caption="Track 0 and the solo PR loop can run in parallel. Multi-agent and the tool builder cannot start in the first 30 days."
  chart="flowchart TB; Start[Day 0] --> City[Track 0: city schema and public reports]; Start --> Agent[MVP 1: solo agent to PR]; City --> GateA[Gate A]; Agent --> GateA; GateA --> Index[MVP 2: indexing and scoped memory]; Index --> Live[MVP 3: bind live events to the city]; Live --> Multi[MVP 4: multi-agent]; Multi --> Tools[MVP 5: tool builder]; Tools --> Learn[MVP 6: bandits on logged trajectories]"
/>

### Track 0 — city foundation and public reports

Build the world schema before the agent is impressive.

Why early: if folders become sectors, files become buildings, tests become a facility, and PRs become the harbor, then events, reports, and later live agents all have a stable target. Integration later is cheaper if the world exists first. It also makes the project demoable before Gate A is green.

The important product decision: **a public report generates a city.**

<Mermaid
  caption="Private in-product world and public report world share a schema. The report world is an immutable snapshot."
  chart="flowchart LR; A[Run finishes or fixture report] --> B[ReportComposer]; B --> C[WorldGenerator]; C --> D[CitySnapshot]; D --> E[Shareable PublicReport URL]; E --> F[City viewer]; E --> G[Panels anchored to buildings]"
/>

A file diff anchors to a building. Tests anchor to the testing facility. A PR anchors to the harbor. Failures put a warning on a district. Agent steps can highlight a path.

Track 0 already runs on a fixture repo. The report route can navigate from sections to file, test, and pull-request anchors. That establishes the city/report contract; it does not establish live task events or autonomous delivery.

### MVP 1 — solo agent to PR, plus trajectories (Gate A)

In parallel: GitHub App, BYOK OpenAI, Daytona sandbox, one coding agent, commit, push, PR.

Until Gate A is green, I am not allowed to ship multi-agent swarms, build the Tool Builder product, turn on active Global Memory, or claim a finished “live autonomous city.”

Allowed: Track 0, fixture report → city, event schema, text inspector, logging memory that is not globally active.

<DataTable
  dataset="gates"
  caption="Fancy layers amplify whatever quality already exists. A tool factory on a bad solo agent is a chaos amplifier."
/>

Gate A, in plain language: a real repo, a real key, a real sandbox, a real PR, 100% of completed runs write a complete trajectory with `reward_v1`, and a real run can mint a PublicReport whose CitySnapshot matches that run.

This is boring on purpose. OpenAI’s Agents SDK direction is also “harness first”: files, shell, sandboxes, skills, compaction — not “spawn a company of agents on day one” ([OpenAI](https://openai.com/index/the-next-evolution-of-the-agents-sdk/)).

## The stack I am starting with, and why

I want to spend engineering time on agent quality, context, memory, and the city — not on inventing infrastructure.

<DataTable
  dataset="stack"
  caption="The current slice is TypeScript, Next.js, React, and Vitest. The remaining rows are architecture choices for later phases, not implemented services."
/>

Rust is not a v1 flex. A typical run is waiting on the model, `npm install`, tests, and git — not on the orchestrator’s allocation speed.

The project card shows the technology set for Tourist as a whole. The current prototype only uses part of it. OpenAI, Postgres, Redis, Pinecone, Qdrant, LangChain, LangGraph, and LangSmith belong to the agent, memory, retrieval, and observability plan. Qdrant is the present vector-search choice; Pinecone is another candidate. Reinforcement learning means a later, evaluated policy over discrete agent decisions, not model-weight training in v1.

BYOK is a product constraint, not a billing footnote. The user’s OpenAI key runs the agent. I meter usage. I do not become an opaque model reseller on day one.

Daytona matters because the agent should outlive the laptop:

```text
Start task → close laptop → sandbox keeps working → tests → PR
```

OpenAI’s sandbox agent model is exactly this shape: the agent is still an agent (instructions, tools, handoffs, guardrails), but the execution boundary is a live sandbox that owns files, commands, and isolation. Providers include Daytona, E2B, Modal, Vercel, Docker, and others ([OpenAI Developers](https://developers.openai.com/api/docs/guides/agents/sandboxes)). I am picking Daytona because the Python SDK path is documented, pause-on-exit preserves the filesystem, and I do not want to own VM lifecycle yet ([Daytona](https://www.daytona.io/docs/en/guides/openai/openai-agents-sdk-with-sandboxes/)).

## Context: copy Cursor’s discovery model, then make retrieval a learned decision

This is the part I am most sure about.

Older agent setups stuffed the prompt with the repo tree, semantic matches, memories, tool catalogs, and chat history “just in case.” Cursor’s current direction is the opposite. Agents start with lightweight facts (OS, git status, recent files) and **discover** the rest through tools ([Katz, 2026](https://cursor.com/blog/dynamic-context-discovery); [Cursor](https://cursor.com/blog/continually-improving-agent-harness)).

The teaching version of that contrast looks like this. These are not Tourist measurements. They are a sketch of why “just in case” is expensive:

<Chart
  dataset="contextComposition"
  type="bar"
  stacked
  title="What actually sits in the prompt"
  xKey="strategy"
  series="contextCompositionSeries"
  yUnit=" tokens"
  caption="Illustrative teaching data, not a Tourist benchmark. Static dump ≈ 48k tokens of 'might help.' Dynamic discovery ≈ 10k tokens of what the agent actually pulled. Cursor's claim is that the second shape can also improve quality by keeping contradictory material out."
/>

### Search is hybrid, not “embed the repo into the prompt”

Cursor reports that semantic search improved coding-agent question accuracy by **12.5% on average** (range 6.5%–23.5% depending on the model), that grep + semantic search together currently perform best, and that the benefit is larger on big codebases. Their A/B tests also showed higher code retention on large repos when semantic search was available ([Heule, Jia & Jain, 2025](https://cursor.com/blog/semsearch)).

<Chart
  dataset="semsearchLift"
  type="bar"
  title="Cursor: semantic search lift on question accuracy"
  xKey="model"
  series="semsearchSeries"
  yUnit="%"
  caption="Source: Cursor Context Bench, reported in 'Improving agent with semantic search' (Nov 2025). These are Cursor's results, not Tourist's. Average +12.5%, range 6.5%–23.5% depending on the model."
/>

<DataTable
  dataset="cursorOnlineAb"
  caption="Cursor's online A/B test, same model, semantic search on vs off. Effect sizes are smaller than the offline accuracy lift because not every agent query needs search."
/>

So the retrieval ladder is:

```text
symbol lookup → ripgrep → LSP → semantic search → memory
```

Not: embed everything and pray.

<Mermaid
  caption="Indexing happens ahead of time. The model does not receive the index. It receives a query interface."
  chart="flowchart LR; A[Tree-sitter chunks] --> B[Symbols]; A --> C[Embeddings in Qdrant]; D[Agent] --> E[Symbol lookup]; E --> F[ripgrep]; F --> G[LSP]; G --> H[Semantic search]; H --> I[Memory]; B --> E; C --> H"
/>

Cursor’s Dropbox writeup is a useful existence proof that indexing hundreds of thousands of files is a product problem, not a prompt problem ([Cursor](https://cursor.com/blog/dropbox)).

### Long tool output becomes a file

`npm test` can be 40k tokens. Truncating it loses the assertion you needed. Cursor writes long tool responses to files and lets the agent `tail` / read slices. They say this reduced unnecessary summarization ([Katz, 2026](https://cursor.com/blog/dynamic-context-discovery)).

Sandbox contract I want:

```json
{
  "exit_code": 1,
  "artifact": "test-output.log",
  "tail": "...AssertionError..."
}
```

Maybe 300 tokens. Then `read_artifact("test-output.log", lines=550:620)` only if needed.

Terminal sessions get the same treatment: addressable artifacts, not an ever-growing transcript welded into context.

### Tools and skills are discovered, not preloaded

This matters even more later, when agents can create tools.

Cursor’s MCP finding: loading every tool description is expensive; syncing descriptions to files and letting the agent look them up cut **total agent tokens by 46.9%** on runs that actually called an MCP tool ([Katz, 2026](https://cursor.com/blog/dynamic-context-discovery)). They also adopted the Agent Skills open standard: a name + description can sit in static context, and the full instructions / scripts are pulled when relevant ([Anthropic](https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills)).

<Chart
  dataset="mcpTokenUsage"
  type="bar"
  title="Cursor: MCP tool descriptions as files"
  xKey="setup"
  series="mcpTokenSeries"
  yUnit=""
  caption="Relative total agent tokens on runs that called an MCP tool, indexed to 100. Cursor reports a 46.9% reduction (statistically significant, with high variance based on the number of MCPs installed). Source: Dynamic context discovery, Jan 2026. Not a Tourist result."
/>

If Tourist ever has hundreds of tools, dumping every schema into the system prompt is how you go broke and also how you confuse the model. The Tool Registry should retrieve a tiny pack per task.

### Summaries must not destroy history

Compaction is lossy. Cursor keeps chat history as files so the agent can recover details the summary dropped.

For Tourist:

- full task history → object storage
- structured events → Postgres
- semantic summaries → Qdrant
- compact summary → current prompt

The agent can `search_task_history("why did we reject Redis?")` instead of trusting a 3k-token amnesia note. That is episodic memory with a retrieval API.

### The extra step: context selection is part of the policy

Cursor is already training retrieval from agent traces: look at what the agent eventually opened, ask a model what should have been retrieved earlier, train embeddings toward that ([Heule, Jia & Jain, 2025](https://cursor.com/blog/semsearch)).

I want that loop, but as a **discrete, logged decision**, not only as a better embedding model.

Every run chooses a `context_budget` (`4k` / `8k` / `16k`), a `tool_pack`, a `memory_pack`. Those choices go into the trajectory. Later, contextual bandits can prefer packs that actually helped the PR get merged, without retraining GPT.

A useful scoring rule for the router:

```text
context_score
= (relevance × expected usefulness × confidence) / token_cost
```

If `auth.ts` is 1,800 tokens at 0.98 relevance and `README.md` is 7,000 tokens at 0.31, the README does not get to eat the window.

<DataTable
  dataset="contextScoreInputs"
  caption="Teaching numbers, not measured retrieval scores. The formula is the point: a long, vaguely relevant file should lose to a short, highly relevant one."
/>

<Chart
  dataset="contextScores"
  type="bar"
  title="Same formula, three files, very different claims on the window"
  xKey="file"
  series="contextScoreSeries"
  caption="Illustrative context_score × 1,000,000. auth.ts wins because it is small and highly relevant. README.md is long and only loosely related, so its score collapses. Start the agent around ~10k tokens of active context even if the available universe is tens of millions."
/>

> The goal is not maximum context. It is minimum sufficient context.

## Memory: persistent, scoped, and hostile to pollution

Memory is how the system stops being session-oriented. It is also how you leak secrets into the next user’s prompt if you are sloppy.

<DataTable
  dataset="memoryScopes"
  caption="Global Memory is the dangerous one. Logging can exist before Gate B. Promotion cannot."
/>

Agents never write `scope=global, status=active` themselves. The pipeline is:

```text
default deny
→ sanitize
→ candidate
→ multi-repo promote
→ budgeted retrieve
→ demote
```

Nothing secret, repo-specific, or user-specific graduates. Active global retrieval stays off until Gate B.

<Mermaid
  caption="No memory should move directly from one conversation into global behaviour."
  chart="flowchart LR; A[Run evidence] --> B[Sanitize]; B --> C[Candidate]; C --> D{Scope?}; D -- User or repo --> E[Scoped store]; D -- Global --> F[Default deny]; F --> G[Multi-repo promote]; G --> H[Budgeted retrieve]; H --> I[Demote if it regresses]; E --> H"
/>

This is less exciting than “the agent remembers everything” and much closer to how you would actually run it.

## Learning: around the model, not of the model

I am going to be strict about language. Tourist does **structured policy learning over trajectories**. It does not “do RL on GPT” in v1.

Non-negotiables:

1. Do not train foundation-model weights in v1.
2. Do not do online PPO against the main LLM.
3. Every completed run writes a trajectory + reward. Incomplete learning records are bugs.
4. Only discrete, versioned decisions are learned.
5. No policy promotion without an eval gate. Logging is not learning.
6. Global Memory graduates under rules.

The supervisor’s action space is small on purpose:

<DataTable
  dataset="supervisorArms"
  caption="This is a contextual bandit problem, not a foundation-model training problem."
/>

The classic result here is that you can learn which arm to pull given context (task type, repo size, languages, CI present, prior failures) without needing a full MDP and without updating GPT. Li et al.’s LinUCB work is the usual citation for contextual bandits with exploration that does not require training a giant policy network ([Li et al., 2010](https://arxiv.org/abs/1003.0146)).

Rewards are delayed and versioned (`reward_v1`): tests, CI, tokens, latency, PR merge/revert, user accept/reject. Retrieval gets credit assignment: if a file was retrieved early and actually used in the successful patch, that retrieval was good. If the agent wandered for 20 searches then found the file, the initial pack was wrong.

Policies go shadow → canary → active. Eval harness before any “self-improving” claim. SWE-bench and related agent benchmarks exist so nobody has to pretend vibes are a metric ([Jimenez et al., 2024](https://www.swebench.com/)). I will still need my own fixture-repo eval (≥ 20 tasks) because Tourist’s loop includes GitHub + sandbox + city reports, not just patch generation.

<Callout title="Why log trajectories from MVP 1 if learning is MVP 6?">
  You cannot reconstruct a faithful trajectory later. If I wait until the bandit code exists, I will have thrown away the only data that could train it.
</Callout>

SWE-agent is useful supporting evidence for investing in the interface, not only the base model: the agent-computer interface changes success rates ([Yang et al., 2024](https://arxiv.org/abs/2405.15793)). Artifacts, tools, and a stop policy are part of that interface.

## The city is the report, then it becomes the runtime

I care about the city more than a dashboard, and less than the PR loop.

Software is spatial whether we admit it or not. You already say “the auth area,” “the payments module,” “the messy legacy folder.” Mapping that to an island is not a gimmick if the mapping is generated from the repo tree and events, not hand-authored.

<DataTable
  dataset="cityMapping"
  caption="The mapping has to be generated. A hand-authored pretty city is a trailer."
/>

<Mermaid
  caption="A file diff anchors to a building. A PR anchors to the harbor. Failures put a warning on a district."
  chart="flowchart TB; Repo[Repository] --> Island[Island]; Folder[Major folder] --> Sector[Sector]; Sub[Subdirectory] --> District[District]; File[File] --> Building[Building]; Tests[Tests] --> Facility[Testing facility]; PR[GitHub PR] --> Harbor[Harbor]; Agent[Agent] --> Character[Character]; Tool[Tool creation] --> Workshop[Workshop]"
/>

Art direction: isometric pixel art, closer to a strategy game than to a Graphviz screenshot. The working viewer uses SVG and DOM components, with one configurable building per file and a generic building when the file type is unmapped. React Three Fiber and Three.js remain possible later choices for a true 3D world; they are not the current renderer.

The building collection is intended to grow toward roughly 75 distinct file-type or semantic-role designs. Equivalent extensions can share a design, while files such as `*.test.ts` can use a testing design instead of an ordinary TypeScript building. Command Center, Research Lab, Workshop, and Harbor are landmarks rather than file buildings. That collection is a target, not a claim that all 75 designs are finished.

Build order for the city:

1. World schema + file-building registry — working first slice
2. Fixture repo → deterministic layout → interactive viewer — working first slice
3. Anchors + fixture PublicReport route — working first slice
4. Persist CitySnapshot
5. Only after Gate A: bind live pawns, construction, harbor traffic

If Gate A regresses, city polish pauses. The report/snapshot path should keep working. A beautiful empty city is a trailer. A shareable report with a generated city for a real run is a product.

## How I will implement it in the repo

Monorepo from the start, because the protocol is the product.

```text
tourist/
├── apps/web                     # city viewer + report pages early
├── apps/api
├── apps/worker
├── services/agent-runtime       # Python
├── services/indexing
├── services/memory
├── services/learning
├── services/report-composer
├── packages/protocol            # World, CitySnapshot, PublicReport, Trajectory, Reward
├── packages/events
├── packages/github
└── packages/world-generator     # Track 0 — not “later”
```

Schema changes update `packages/protocol` in the same PR. If the city and the agent disagree about what a “building” is, that is a protocol bug, not a frontend bug.

Branching follows the gates: city foundation and fixture reports, cloud agent, trajectories, Gate A, then indexing/memory, living city, multi-agent, tools, learning. Track 0 and Phase 1 can progress in parallel, with the completed city slice providing a concrete report contract for the agent work.

## Original first-month plan

<DataTable
  dataset="firstThirtyDays"
  caption="The original sequence, retained as a plan rather than a completed milestone. The fixture city and report route are running; the real solo-agent PR and trajectory gates remain open."
/>

The fixture PublicReport and its navigable city now cover the first part of that contract. Gate A still needs:

- a solo agent that can open a PR on a real repo
- every completed run leaving a trajectory behind
- a real run producing a PublicReport whose CitySnapshot matches that run

## What I am afraid of

<DataTable
  dataset="fears"
  caption="Gates exist because I do not trust future-me to stay boring."
/>

The honest bet: **harness + discovery + persistence + a spatial report** is still underbuilt relative to “wrap the latest model in a chat UI.” Cursor is iterating on the harness in public. OpenAI is standardizing sandbox agents and skills. The open space I want is a cloud loop that remembers, learns discrete policies from PR outcomes, and makes the work visible as a city you can share.

<MarginNote>
  A log stream tells you the agent did something. A city tells you where.
</MarginNote>

If that city is only a skin, the project failed. If the city is how a run becomes understandable — files, tests, harbor, failures — then the visualization is doing engineering work, not marketing.

That is the plan. The fixture city exists. Next the cloud loop has to earn the harbor: a real sandbox run, a tested change, a PR, and a trajectory that lets us explain what happened.

The next time a cloud agent opens a PR while your laptop is closed, the interesting question is not only whether the patch is right.

It is what the system remembered from last Tuesday, what it refused to load into the window, and whether you can point at a building and see the work.

<References title="Sources and further reading" />

## Companion structured data

```json
{
  "dataNote": "Cursor figures are from published Cursor research. Token composition and context scores are teaching illustrations, not Tourist measurements. Tourist's city foundation is runnable; its cloud agent loop has not shipped.",
  "mcpTokenSeries": [
    {
      "key": "tokens",
      "label": "Relative total agent tokens",
      "color": "#b65332"
    }
  ],
  "mcpTokenUsage": [
    {
      "setup": "MCP tools always loaded",
      "tokens": 100
    },
    {
      "setup": "Tool descriptions as files",
      "tokens": 53.1
    }
  ],
  "semsearchSeries": [
    {
      "key": "lift",
      "label": "Question accuracy lift",
      "color": "#356b51"
    }
  ],
  "semsearchLift": [
    {
      "model": "Lowest model in Cursor's set",
      "lift": 6.5
    },
    {
      "model": "Average across models",
      "lift": 12.5
    },
    {
      "model": "Highest model in Cursor's set",
      "lift": 23.5
    }
  ],
  "contextCompositionSeries": [
    {
      "key": "statusAndTree",
      "label": "Repo status / tree",
      "color": "#356b51"
    },
    {
      "key": "retrievedCode",
      "label": "Retrieved code",
      "color": "#4a8c6a"
    },
    {
      "key": "memories",
      "label": "Memories",
      "color": "#88619a"
    },
    {
      "key": "toolSchemas",
      "label": "Tool schemas",
      "color": "#b65332"
    },
    {
      "key": "history",
      "label": "History / summaries",
      "color": "#c4893a"
    }
  ],
  "contextComposition": [
    {
      "strategy": "Static dump",
      "statusAndTree": 12000,
      "retrievedCode": 8000,
      "memories": 5000,
      "toolSchemas": 14000,
      "history": 9000
    },
    {
      "strategy": "Dynamic discovery",
      "statusAndTree": 300,
      "retrievedCode": 3500,
      "memories": 800,
      "toolSchemas": 900,
      "history": 4500
    }
  ],
  "contextScoreSeries": [
    {
      "key": "score",
      "label": "Illustrative context score",
      "color": "#b65332"
    }
  ],
  "contextScores": [
    {
      "file": "auth.ts",
      "score": 416.5
    },
    {
      "file": "session.ts",
      "score": 146.7
    },
    {
      "file": "README.md",
      "score": 10.6
    }
  ],
  "contextScoreInputs": [
    {
      "File": "auth.ts",
      "Tokens": "1,800",
      "Relevance": "0.98",
      "Expected usefulness": "0.90",
      "Confidence": "0.85",
      "Score × 1,000,000": "416.5"
    },
    {
      "File": "session.ts",
      "Tokens": "2,400",
      "Relevance": "0.71",
      "Expected usefulness": "0.62",
      "Confidence": "0.80",
      "Score × 1,000,000": "146.7"
    },
    {
      "File": "README.md",
      "Tokens": "7,000",
      "Relevance": "0.31",
      "Expected usefulness": "0.40",
      "Confidence": "0.60",
      "Score × 1,000,000": "10.6"
    }
  ],
  "cursorOnlineAb": [
    {
      "Metric": "Code retention, all agent queries",
      "Reported change": "+0.3% with semantic search",
      "Caveat": "Many queries do not need search"
    },
    {
      "Metric": "Code retention, repos with 1,000+ files",
      "Reported change": "+2.6% with semantic search",
      "Caveat": "Larger effect on large codebases"
    },
    {
      "Metric": "Dissatisfied follow-up requests",
      "Reported change": "+2.2% when semantic search was removed",
      "Caveat": "Measured across all agent queries"
    }
  ],
  "whatTouristIs": [
    {
      "Layer": "Context engine",
      "Job": "Start light; pull files, logs, skills, and memories on demand",
      "Not this": "Paste the repo tree, tool catalog, and chat history into every prompt"
    },
    {
      "Layer": "Scoped memory",
      "Job": "Remember user, repo, and task lessons with promotion rules",
      "Not this": "A global junk drawer the agent can write into freely"
    },
    {
      "Layer": "Trajectories and rewards",
      "Job": "Log every completed run with reward_v1 from day one",
      "Not this": "Train GPT weights, or reconstruct traces later from vibes"
    },
    {
      "Layer": "Generated city",
      "Job": "Turn a run into a shareable spatial report, then later a live world",
      "Not this": "A dashboard skin, or a claim that the city is already autonomous"
    }
  ],
  "stack": [
    {
      "Layer": "City foundation (current)",
      "Choice": "Next.js, React, TypeScript, SVG/DOM, Vitest",
      "Why": "Versioned protocol, deterministic world generator, per-file components, fixture report, and tests"
    },
    {
      "Layer": "3D world (possible later)",
      "Choice": "React Three Fiber + Three.js",
      "Why": "An option if the live world needs true 3D; the working viewer is isometric SVG/DOM"
    },
    {
      "Layer": "API (planned)",
      "Choice": "Fastify + TypeScript",
      "Why": "Control plane for GitHub, events, and auth"
    },
    {
      "Layer": "Agent runtime (planned)",
      "Choice": "Python + OpenAI Agents SDK",
      "Why": "Python-first SDK, Responses API, native sandbox execution"
    },
    {
      "Layer": "Sandboxes (planned)",
      "Choice": "Daytona",
      "Why": "Isolated cloud workspaces, pause/resume, first-class Agents SDK client"
    },
    {
      "Layer": "GitHub delivery (planned)",
      "Choice": "GitHub App + Octokit",
      "Why": "PRs are the delivery primitive, not chat transcripts"
    },
    {
      "Layer": "Primary DB (planned)",
      "Choice": "Postgres (Neon)",
      "Why": "Source of truth for tasks, trajectories, memory records, reports"
    },
    {
      "Layer": "Vectors (planned)",
      "Choice": "Qdrant",
      "Why": "Hybrid search for code chunks and memories"
    },
    {
      "Layer": "Retrieval alternative (planned)",
      "Choice": "Pinecone",
      "Why": "A candidate to compare with Qdrant when the indexing and retrieval path exists"
    },
    {
      "Layer": "Agent tooling (planned)",
      "Choice": "LangChain, LangGraph, LangSmith",
      "Why": "Possible orchestration and tracing tools; no dependency on them in the current city prototype"
    },
    {
      "Layer": "Cache / jobs (planned)",
      "Choice": "Redis + BullMQ",
      "Why": "Hot state and background work; Temporal later if workflows get gnarly"
    },
    {
      "Layer": "Artifacts (planned)",
      "Choice": "R2 / S3",
      "Why": "Trajectories, logs, CitySnapshots"
    },
    {
      "Layer": "Indexing (planned)",
      "Choice": "Tree-sitter, ripgrep, LSP, embeddings",
      "Why": "Deterministic search first; semantic search as a complement"
    }
  ],
  "memoryScopes": [
    {
      "Scope": "User",
      "Holds": "Preferences such as TypeScript, smaller PRs, pnpm",
      "Write rule": "User- or workspace-scoped; never promoted globally"
    },
    {
      "Scope": "Codebase",
      "Holds": "Architecture, conventions, fragile areas, past agent work on this repo",
      "Write rule": "Repo-scoped; versioned with the tree"
    },
    {
      "Scope": "Episodic",
      "Holds": "This task failed because X, then Y worked",
      "Write rule": "Task-scoped; retrievable, not always in the prompt"
    },
    {
      "Scope": "Procedural / tool",
      "Holds": "For this class of problem, use that tool",
      "Write rule": "Candidate until replayed"
    },
    {
      "Scope": "Shared task",
      "Holds": "Scratch space for later multi-agent runs",
      "Write rule": "Exists in the schema; unused until Gate C"
    },
    {
      "Scope": "Global",
      "Holds": "Reusable patterns across repos",
      "Write rule": "Default deny. Agents never write scope=global, status=active"
    }
  ],
  "supervisorArms": [
    {
      "Decision": "model",
      "Example arms": "fast / coding / reasoning"
    },
    {
      "Decision": "topology",
      "Example arms": "solo_coder / coder_tester / …"
    },
    {
      "Decision": "context_budget",
      "Example arms": "4k / 8k / 16k"
    },
    {
      "Decision": "tool_pack",
      "Example arms": "retrieved tool ids"
    },
    {
      "Decision": "memory_pack",
      "Example arms": "retrieved memory ids"
    },
    {
      "Decision": "stop_policy",
      "Example arms": "stop_on_green / max_3_retries / ask human"
    },
    {
      "Decision": "test_policy",
      "Example arms": "unit_only / unit+lint / full CI"
    }
  ],
  "gates": [
    {
      "Gate": "A / MVP 1",
      "What must be true": "Real repo, BYOK, Daytona sandbox, real PR, complete trajectory with reward_v1, PublicReport whose CitySnapshot matches the run",
      "Still forbidden": "Swarms, Tool Builder product, active Global Memory, live-autonomous-city claims"
    },
    {
      "Gate": "B / MVP 2",
      "What must be true": "Indexing plus scoped memory; Global Memory only after sanitization and promotion",
      "Still forbidden": "Letting the agent self-promote global rules"
    },
    {
      "Gate": "C / MVP 3–4",
      "What must be true": "Live agent events bind onto the Track 0 city, then multi-agent",
      "Still forbidden": "Starting multi-agent in the first 30 days"
    },
    {
      "Gate": "D / MVP 5–6",
      "What must be true": "Tool builder, then bandits / retrieval learning behind an eval gate",
      "Still forbidden": "Calling logging 'learning', or promoting a policy on vibes"
    }
  ],
  "cityMapping": [
    {
      "Software": "Repository",
      "City object": "Island / city"
    },
    {
      "Software": "Major folder",
      "City object": "Sector"
    },
    {
      "Software": "Subdirectory",
      "City object": "District"
    },
    {
      "Software": "File",
      "City object": "Building"
    },
    {
      "Software": "Agent",
      "City object": "Character"
    },
    {
      "Software": "Tool creation",
      "City object": "Workshop"
    },
    {
      "Software": "Tests",
      "City object": "Testing facility"
    },
    {
      "Software": "GitHub / PRs",
      "City object": "Harbor"
    }
  ],
  "firstThirtyDays": [
    {
      "Week": "1",
      "City / reports": "Scaffold, world schema, art spikes",
      "Agent": "GitHub App skeleton"
    },
    {
      "Week": "2",
      "City / reports": "Fixture repo renders as a city",
      "Agent": "Daytona hello-world + BYOK path"
    },
    {
      "Week": "3",
      "City / reports": "PublicReport anchors + share URL stub",
      "Agent": "Edit → commit → push → PR"
    },
    {
      "Week": "4",
      "City / reports": "CitySnapshot persist; polish report page",
      "Agent": "Trajectory + reward_v1; one real run → report → city"
    }
  ],
  "fears": [
    {
      "Fear": "Six products at once",
      "Control": "Quality gates. Fancy layers cannot start until Gate A is green"
    },
    {
      "Fear": "Vague 'RL'",
      "Control": "A small, versioned decision table. No PPO against GPT in v1"
    },
    {
      "Fear": "Global memory as a secret-shaped junk drawer",
      "Control": "Default deny, sanitization, promotion, demotion"
    },
    {
      "Fear": "City polish replacing agent quality",
      "Control": "Track 0 is schema, viewer, and reports. Gate A still blocks live claims"
    },
    {
      "Fear": "Multi-agent too early",
      "Control": "One agent that ships beats a planner-coder-reviewer triangle that argues in a sandbox"
    }
  ],
  "citations": [
    {
      "title": "Dynamic context discovery",
      "publisher": "Jediah Katz, Cursor Blog, 6 Jan 2026",
      "url": "https://cursor.com/blog/dynamic-context-discovery",
      "note": "Static vs dynamic context; long tool outputs as files; chat history as files after summarization; Agent Skills; MCP tool descriptions as files; 46.9% fewer total agent tokens on MCP-calling runs; terminals as files."
    },
    {
      "title": "Continually improving our agent harness",
      "publisher": "Cursor",
      "url": "https://cursor.com/blog/continually-improving-agent-harness",
      "note": "Less static preload; agents start light and discover."
    },
    {
      "title": "Improving agent with semantic search",
      "publisher": "Stefan Heule, Emily Jia & Naman Jain, Cursor Blog, 6 Nov 2025",
      "url": "https://cursor.com/blog/semsearch",
      "note": "+12.5% average question accuracy with semantic search (range 6.5%–23.5%); grep + semantic together best; trace-trained embeddings; code retention and follow-up A/B results."
    },
    {
      "title": "Dropbox uses Cursor to index over 550,000 files",
      "publisher": "Cursor",
      "url": "https://cursor.com/blog/dropbox",
      "note": "Indexing at scale is a product problem. The model does not eat the whole corpus."
    },
    {
      "title": "The next evolution of the Agents SDK",
      "publisher": "OpenAI",
      "url": "https://openai.com/index/the-next-evolution-of-the-agents-sdk/",
      "note": "Model-native harness, sandbox execution, skills / AGENTS.md / shell / apply-patch, providers including Daytona."
    },
    {
      "title": "Sandbox agents",
      "publisher": "OpenAI Developers",
      "url": "https://developers.openai.com/api/docs/guides/agents/sandboxes",
      "note": "SandboxAgent still has tools, handoffs, and guardrails; the execution boundary is the sandbox."
    },
    {
      "title": "Using the OpenAI Agents SDK with Daytona Sandboxes",
      "publisher": "Daytona",
      "url": "https://www.daytona.io/docs/en/guides/openai/openai-agents-sdk-with-sandboxes/",
      "note": "Shell and filesystem in-sandbox; pause_on_exit / resume; isolated cloud execution."
    },
    {
      "title": "Equipping agents for the real world with Agent Skills",
      "publisher": "Anthropic",
      "url": "https://www.anthropic.com/engineering/equipping-agents-for-the-real-world-with-agent-skills",
      "note": "Skills as files with a name and description; load full instructions on demand."
    },
    {
      "title": "A Contextual-Bandit Approach to Personalized News Article Recommendation",
      "publisher": "Li, Chu, Langford, Schapire, WWW 2010",
      "url": "https://arxiv.org/abs/1003.0146",
      "note": "LinUCB: discrete arms plus context, without training the underlying model."
    },
    {
      "title": "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?",
      "publisher": "Jimenez et al., ICLR 2024",
      "url": "https://www.swebench.com/",
      "note": "Opening a PR is not enough. Held-out eval still matters."
    },
    {
      "title": "SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering",
      "publisher": "Yang et al., NeurIPS 2024",
      "url": "https://arxiv.org/abs/2405.15793",
      "note": "Interface and harness design change agent success, not only the base model."
    },
    {
      "title": "React Three Fiber documentation",
      "publisher": "Poimandres",
      "url": "https://docs.pmnd.rs/react-three-fiber",
      "note": "React renderer for Three.js — the city viewer stack."
    },
    {
      "title": "GitHub Apps documentation",
      "publisher": "GitHub",
      "url": "https://docs.github.com/en/apps",
      "note": "Installation, permissions, PR creation — delivery is GitHub-native."
    },
    {
      "title": "Model Context Protocol",
      "publisher": "Anthropic / MCP",
      "url": "https://modelcontextprotocol.io/",
      "note": "Why tool catalogs explode; background for Cursor's MCP token result."
    }
  ]
}
```
