How I'm Building Tourist: A Cloud Coding Agent That Lives in a City
Tourist's city foundation is running. Here is the build plan for the cloud coding agent, scoped memory, dynamic context, GitHub delivery, and learning that should follow.

Most coding agents still behave like a very smart intern who forgets the office every morning.
You describe a bug. The agent searches, finds auth.ts, patches it, runs tests, and opens a pull request. Tomorrow you describe a related bug. It searches the same folders, rereads the same README, retries the approach that already failed on Tuesday, and dumps half the repo into the prompt “just in case.”
The model is not the bottleneck anymore. The session is.
I usually write about systems I have already poked at. Tourist now has a working city foundation. The larger product I am building looks like this:
Connect a GitHub repo
→ bring your own OpenAI key
→ describe the work
→ a cloud agent inspects the codebase, edits and tests in an isolated sandbox
→ a pull request shows up
The longer version is the actual product. Most agents are still chat sessions with tools. Tourist is an attempt to make something closer to a persistent engineering environment: memory across tasks, context that is discovered instead of dumped, learning from outcomes, and a visual world that makes autonomous work inspectable.
This post is the build plan and a progress marker. The first city and report slice works. The cloud agent loop, persistent memory, tool creation, and learning are still work ahead.
The problem I am actually trying to solve
Coding agents got good fast. They can search a repo, edit files, run tests, and open PRs. What is still weak is everything around the session:
- they rediscover the same architecture every time
- they forget the approach that already failed last Tuesday
- they preload too much context, then drown in it
- they have a fixed toolbelt
- they do not systematically get better from merged PRs, reverted PRs, flaky tests, or “please don’t do it that way”
- when they work in the cloud, you get a log stream, not a place you can look at
I have already written about two pieces of this: memory as a coordination layer, and agents that know when to stop. Tourist is the project that would have to survive those constraints in production, then make the work visible.
Cursor’s own writing on this is unusually clear. In Dynamic context discovery, they argue that as models get better as agents, you should give them fewer details up front and make it cheap to pull the rest. Their contrast is static context versus dynamic context discovery. The claim is not just “saves tokens.” They say it can also improve quality by keeping contradictory or irrelevant material out of the window (Katz, 2026).
That is the first architectural tenet I am copying:
Huge available context. Tiny active context.
The rest of Tourist is what happens if you take that seriously, then add persistence, GitHub-native delivery, and a city as the reporting surface.
What Tourist is (and is not)
The user flow for v1 is boring on purpose. Connect GitHub, paste an OpenAI key, describe a task, wait. The interesting work is under that loop.
Under that loop I am building four things that most chat-wrapped agents skip:
| Layer | Job | Not this |
|---|---|---|
| Context engine | Start light; pull files, logs, skills, and memories on demand | Paste the repo tree, tool catalog, and chat history into every prompt |
| Scoped memory | Remember user, repo, and task lessons with promotion rules | A global junk drawer the agent can write into freely |
| Trajectories and rewards | Log every completed run with reward_v1 from day one | Train GPT weights, or reconstruct traces later from vibes |
| Generated city | Turn a run into a shareable spatial report, then later a live world | A dashboard skin, or a claim that the city is already autonomous |
What I am not building, at least not first:
- a new foundation model
- online PPO against GPT
- Kubernetes as a personality trait
- a custom sandbox runtime
- a 40-agent swarm
- a VS Code clone
- a claim that “the city is alive” before a solo agent can open a real PR
If I confuse those, I will spend three months making buildings prettier while the agent still cannot ship.
Two tracks, not six products
The failure mode for a project like this is obvious. Memory, multi-agent, tool factories, RL, and a 3D world are five startups. I wrote quality gates into the PRD so I cannot “just quickly add a swarm.”
There are two tracks. The first Track 0 slice is runnable; the solo PR loop remains the next major gate. Everything else is gated.
Track 0 — city foundation and public reports
Build the world schema before the agent is impressive.
Why early: if folders become sectors, files become buildings, tests become a facility, and PRs become the harbor, then events, reports, and later live agents all have a stable target. Integration later is cheaper if the world exists first. It also makes the project demoable before Gate A is green.
The important product decision: a public report generates a city.
A file diff anchors to a building. Tests anchor to the testing facility. A PR anchors to the harbor. Failures put a warning on a district. Agent steps can highlight a path.
Track 0 already runs on a fixture repo. The report route can navigate from sections to file, test, and pull-request anchors. That establishes the city/report contract; it does not establish live task events or autonomous delivery.
MVP 1 — solo agent to PR, plus trajectories (Gate A)
In parallel: GitHub App, BYOK OpenAI, Daytona sandbox, one coding agent, commit, push, PR.
Until Gate A is green, I am not allowed to ship multi-agent swarms, build the Tool Builder product, turn on active Global Memory, or claim a finished “live autonomous city.”
Allowed: Track 0, fixture report → city, event schema, text inspector, logging memory that is not globally active.
| Gate | What must be true | Still forbidden |
|---|---|---|
| A / MVP 1 | Real repo, BYOK, Daytona sandbox, real PR, complete trajectory with reward_v1, PublicReport whose CitySnapshot matches the run | Swarms, Tool Builder product, active Global Memory, live-autonomous-city claims |
| B / MVP 2 | Indexing plus scoped memory; Global Memory only after sanitization and promotion | Letting the agent self-promote global rules |
| C / MVP 3–4 | Live agent events bind onto the Track 0 city, then multi-agent | Starting multi-agent in the first 30 days |
| D / MVP 5–6 | Tool builder, then bandits / retrieval learning behind an eval gate | Calling logging 'learning', or promoting a policy on vibes |
Gate A, in plain language: a real repo, a real key, a real sandbox, a real PR, 100% of completed runs write a complete trajectory with reward_v1, and a real run can mint a PublicReport whose CitySnapshot matches that run.
This is boring on purpose. OpenAI’s Agents SDK direction is also “harness first”: files, shell, sandboxes, skills, compaction — not “spawn a company of agents on day one” (OpenAI).
The stack I am starting with, and why
I want to spend engineering time on agent quality, context, memory, and the city — not on inventing infrastructure.
| Layer | Choice | Why |
|---|---|---|
| City foundation (current) | Next.js, React, TypeScript, SVG/DOM, Vitest | Versioned protocol, deterministic world generator, per-file components, fixture report, and tests |
| 3D world (possible later) | React Three Fiber + Three.js | An option if the live world needs true 3D; the working viewer is isometric SVG/DOM |
| API (planned) | Fastify + TypeScript | Control plane for GitHub, events, and auth |
| Agent runtime (planned) | Python + OpenAI Agents SDK | Python-first SDK, Responses API, native sandbox execution |
| Sandboxes (planned) | Daytona | Isolated cloud workspaces, pause/resume, first-class Agents SDK client |
| GitHub delivery (planned) | GitHub App + Octokit | PRs are the delivery primitive, not chat transcripts |
| Primary DB (planned) | Postgres (Neon) | Source of truth for tasks, trajectories, memory records, reports |
| Vectors (planned) | Qdrant | Hybrid search for code chunks and memories |
| Retrieval alternative (planned) | Pinecone | A candidate to compare with Qdrant when the indexing and retrieval path exists |
| Agent tooling (planned) | LangChain, LangGraph, LangSmith | Possible orchestration and tracing tools; no dependency on them in the current city prototype |
| Cache / jobs (planned) | Redis + BullMQ | Hot state and background work; Temporal later if workflows get gnarly |
| Artifacts (planned) | R2 / S3 | Trajectories, logs, CitySnapshots |
| Indexing (planned) | Tree-sitter, ripgrep, LSP, embeddings | Deterministic search first; semantic search as a complement |
Rust is not a v1 flex. A typical run is waiting on the model, npm install, tests, and git — not on the orchestrator’s allocation speed.
The project card shows the technology set for Tourist as a whole. The current prototype only uses part of it. OpenAI, Postgres, Redis, Pinecone, Qdrant, LangChain, LangGraph, and LangSmith belong to the agent, memory, retrieval, and observability plan. Qdrant is the present vector-search choice; Pinecone is another candidate. Reinforcement learning means a later, evaluated policy over discrete agent decisions, not model-weight training in v1.
BYOK is a product constraint, not a billing footnote. The user’s OpenAI key runs the agent. I meter usage. I do not become an opaque model reseller on day one.
Daytona matters because the agent should outlive the laptop:
Start task → close laptop → sandbox keeps working → tests → PR
OpenAI’s sandbox agent model is exactly this shape: the agent is still an agent (instructions, tools, handoffs, guardrails), but the execution boundary is a live sandbox that owns files, commands, and isolation. Providers include Daytona, E2B, Modal, Vercel, Docker, and others (OpenAI Developers). I am picking Daytona because the Python SDK path is documented, pause-on-exit preserves the filesystem, and I do not want to own VM lifecycle yet (Daytona).
Context: copy Cursor’s discovery model, then make retrieval a learned decision
This is the part I am most sure about.
Older agent setups stuffed the prompt with the repo tree, semantic matches, memories, tool catalogs, and chat history “just in case.” Cursor’s current direction is the opposite. Agents start with lightweight facts (OS, git status, recent files) and discover the rest through tools (Katz, 2026; Cursor).
The teaching version of that contrast looks like this. These are not Tourist measurements. They are a sketch of why “just in case” is expensive:
Loading chart… Data is available below.
View chart data
| strategy | Repo status / tree ( tokens) | Retrieved code ( tokens) | Memories ( tokens) | Tool schemas ( tokens) | History / summaries ( tokens) |
|---|---|---|---|---|---|
| Static dump | 12000 | 8000 | 5000 | 14000 | 9000 |
| Dynamic discovery | 300 | 3500 | 800 | 900 | 4500 |
Search is hybrid, not “embed the repo into the prompt”
Cursor reports that semantic search improved coding-agent question accuracy by 12.5% on average (range 6.5%–23.5% depending on the model), that grep + semantic search together currently perform best, and that the benefit is larger on big codebases. Their A/B tests also showed higher code retention on large repos when semantic search was available (Heule, Jia & Jain, 2025).
Loading chart… Data is available below.
View chart data
| model | Question accuracy lift (%) |
|---|---|
| Lowest model in Cursor's set | 6.5 |
| Average across models | 12.5 |
| Highest model in Cursor's set | 23.5 |
| Metric | Reported change | Caveat |
|---|---|---|
| Code retention, all agent queries | +0.3% with semantic search | Many queries do not need search |
| Code retention, repos with 1,000+ files | +2.6% with semantic search | Larger effect on large codebases |
| Dissatisfied follow-up requests | +2.2% when semantic search was removed | Measured across all agent queries |
So the retrieval ladder is:
symbol lookup → ripgrep → LSP → semantic search → memory
Not: embed everything and pray.
Cursor’s Dropbox writeup is a useful existence proof that indexing hundreds of thousands of files is a product problem, not a prompt problem (Cursor).
Long tool output becomes a file
npm test can be 40k tokens. Truncating it loses the assertion you needed. Cursor writes long tool responses to files and lets the agent tail / read slices. They say this reduced unnecessary summarization (Katz, 2026).
Sandbox contract I want:
{
"exit_code": 1,
"artifact": "test-output.log",
"tail": "...AssertionError..."
}
Maybe 300 tokens. Then read_artifact("test-output.log", lines=550:620) only if needed.
Terminal sessions get the same treatment: addressable artifacts, not an ever-growing transcript welded into context.
Tools and skills are discovered, not preloaded
This matters even more later, when agents can create tools.
Cursor’s MCP finding: loading every tool description is expensive; syncing descriptions to files and letting the agent look them up cut total agent tokens by 46.9% on runs that actually called an MCP tool (Katz, 2026). They also adopted the Agent Skills open standard: a name + description can sit in static context, and the full instructions / scripts are pulled when relevant (Anthropic).
Loading chart… Data is available below.
View chart data
| setup | Relative total agent tokens |
|---|---|
| MCP tools always loaded | 100 |
| Tool descriptions as files | 53.1 |
If Tourist ever has hundreds of tools, dumping every schema into the system prompt is how you go broke and also how you confuse the model. The Tool Registry should retrieve a tiny pack per task.
Summaries must not destroy history
Compaction is lossy. Cursor keeps chat history as files so the agent can recover details the summary dropped.
For Tourist:
- full task history → object storage
- structured events → Postgres
- semantic summaries → Qdrant
- compact summary → current prompt
The agent can search_task_history("why did we reject Redis?") instead of trusting a 3k-token amnesia note. That is episodic memory with a retrieval API.
The extra step: context selection is part of the policy
Cursor is already training retrieval from agent traces: look at what the agent eventually opened, ask a model what should have been retrieved earlier, train embeddings toward that (Heule, Jia & Jain, 2025).
I want that loop, but as a discrete, logged decision, not only as a better embedding model.
Every run chooses a context_budget (4k / 8k / 16k), a tool_pack, a memory_pack. Those choices go into the trajectory. Later, contextual bandits can prefer packs that actually helped the PR get merged, without retraining GPT.
A useful scoring rule for the router:
context_score
= (relevance × expected usefulness × confidence) / token_cost
If auth.ts is 1,800 tokens at 0.98 relevance and README.md is 7,000 tokens at 0.31, the README does not get to eat the window.
| File | Tokens | Relevance | Expected usefulness | Confidence | Score × 1,000,000 |
|---|---|---|---|---|---|
| auth.ts | 1,800 | 0.98 | 0.90 | 0.85 | 416.5 |
| session.ts | 2,400 | 0.71 | 0.62 | 0.80 | 146.7 |
| README.md | 7,000 | 0.31 | 0.40 | 0.60 | 10.6 |
Loading chart… Data is available below.
View chart data
| file | Illustrative context score |
|---|---|
| auth.ts | 416.5 |
| session.ts | 146.7 |
| README.md | 10.6 |
The goal is not maximum context. It is minimum sufficient context.
Memory: persistent, scoped, and hostile to pollution
Memory is how the system stops being session-oriented. It is also how you leak secrets into the next user’s prompt if you are sloppy.
| Scope | Holds | Write rule |
|---|---|---|
| User | Preferences such as TypeScript, smaller PRs, pnpm | User- or workspace-scoped; never promoted globally |
| Codebase | Architecture, conventions, fragile areas, past agent work on this repo | Repo-scoped; versioned with the tree |
| Episodic | This task failed because X, then Y worked | Task-scoped; retrievable, not always in the prompt |
| Procedural / tool | For this class of problem, use that tool | Candidate until replayed |
| Shared task | Scratch space for later multi-agent runs | Exists in the schema; unused until Gate C |
| Global | Reusable patterns across repos | Default deny. Agents never write scope=global, status=active |
Agents never write scope=global, status=active themselves. The pipeline is:
default deny
→ sanitize
→ candidate
→ multi-repo promote
→ budgeted retrieve
→ demote
Nothing secret, repo-specific, or user-specific graduates. Active global retrieval stays off until Gate B.
This is less exciting than “the agent remembers everything” and much closer to how you would actually run it.
Learning: around the model, not of the model
I am going to be strict about language. Tourist does structured policy learning over trajectories. It does not “do RL on GPT” in v1.
Non-negotiables:
- Do not train foundation-model weights in v1.
- Do not do online PPO against the main LLM.
- Every completed run writes a trajectory + reward. Incomplete learning records are bugs.
- Only discrete, versioned decisions are learned.
- No policy promotion without an eval gate. Logging is not learning.
- Global Memory graduates under rules.
The supervisor’s action space is small on purpose:
| Decision | Example arms |
|---|---|
| model | fast / coding / reasoning |
| topology | solo_coder / coder_tester / … |
| context_budget | 4k / 8k / 16k |
| tool_pack | retrieved tool ids |
| memory_pack | retrieved memory ids |
| stop_policy | stop_on_green / max_3_retries / ask human |
| test_policy | unit_only / unit+lint / full CI |
The classic result here is that you can learn which arm to pull given context (task type, repo size, languages, CI present, prior failures) without needing a full MDP and without updating GPT. Li et al.’s LinUCB work is the usual citation for contextual bandits with exploration that does not require training a giant policy network (Li et al., 2010).
Rewards are delayed and versioned (reward_v1): tests, CI, tokens, latency, PR merge/revert, user accept/reject. Retrieval gets credit assignment: if a file was retrieved early and actually used in the successful patch, that retrieval was good. If the agent wandered for 20 searches then found the file, the initial pack was wrong.
Policies go shadow → canary → active. Eval harness before any “self-improving” claim. SWE-bench and related agent benchmarks exist so nobody has to pretend vibes are a metric (Jimenez et al., 2024). I will still need my own fixture-repo eval (≥ 20 tasks) because Tourist’s loop includes GitHub + sandbox + city reports, not just patch generation.
SWE-agent is useful supporting evidence for investing in the interface, not only the base model: the agent-computer interface changes success rates (Yang et al., 2024). Artifacts, tools, and a stop policy are part of that interface.
The city is the report, then it becomes the runtime
I care about the city more than a dashboard, and less than the PR loop.
Software is spatial whether we admit it or not. You already say “the auth area,” “the payments module,” “the messy legacy folder.” Mapping that to an island is not a gimmick if the mapping is generated from the repo tree and events, not hand-authored.
| Software | City object |
|---|---|
| Repository | Island / city |
| Major folder | Sector |
| Subdirectory | District |
| File | Building |
| Agent | Character |
| Tool creation | Workshop |
| Tests | Testing facility |
| GitHub / PRs | Harbor |
Art direction: isometric pixel art, closer to a strategy game than to a Graphviz screenshot. The working viewer uses SVG and DOM components, with one configurable building per file and a generic building when the file type is unmapped. React Three Fiber and Three.js remain possible later choices for a true 3D world; they are not the current renderer.
The building collection is intended to grow toward roughly 75 distinct file-type or semantic-role designs. Equivalent extensions can share a design, while files such as *.test.ts can use a testing design instead of an ordinary TypeScript building. Command Center, Research Lab, Workshop, and Harbor are landmarks rather than file buildings. That collection is a target, not a claim that all 75 designs are finished.
Build order for the city:
- World schema + file-building registry — working first slice
- Fixture repo → deterministic layout → interactive viewer — working first slice
- Anchors + fixture PublicReport route — working first slice
- Persist CitySnapshot
- Only after Gate A: bind live pawns, construction, harbor traffic
If Gate A regresses, city polish pauses. The report/snapshot path should keep working. A beautiful empty city is a trailer. A shareable report with a generated city for a real run is a product.
How I will implement it in the repo
Monorepo from the start, because the protocol is the product.
tourist/
├── apps/web # city viewer + report pages early
├── apps/api
├── apps/worker
├── services/agent-runtime # Python
├── services/indexing
├── services/memory
├── services/learning
├── services/report-composer
├── packages/protocol # World, CitySnapshot, PublicReport, Trajectory, Reward
├── packages/events
├── packages/github
└── packages/world-generator # Track 0 — not “later”
Schema changes update packages/protocol in the same PR. If the city and the agent disagree about what a “building” is, that is a protocol bug, not a frontend bug.
Branching follows the gates: city foundation and fixture reports, cloud agent, trajectories, Gate A, then indexing/memory, living city, multi-agent, tools, learning. Track 0 and Phase 1 can progress in parallel, with the completed city slice providing a concrete report contract for the agent work.
Original first-month plan
| Week | City / reports | Agent |
|---|---|---|
| 1 | Scaffold, world schema, art spikes | GitHub App skeleton |
| 2 | Fixture repo renders as a city | Daytona hello-world + BYOK path |
| 3 | PublicReport anchors + share URL stub | Edit → commit → push → PR |
| 4 | CitySnapshot persist; polish report page | Trajectory + reward_v1; one real run → report → city |
The fixture PublicReport and its navigable city now cover the first part of that contract. Gate A still needs:
- a solo agent that can open a PR on a real repo
- every completed run leaving a trajectory behind
- a real run producing a PublicReport whose CitySnapshot matches that run
What I am afraid of
| Fear | Control |
|---|---|
| Six products at once | Quality gates. Fancy layers cannot start until Gate A is green |
| Vague 'RL' | A small, versioned decision table. No PPO against GPT in v1 |
| Global memory as a secret-shaped junk drawer | Default deny, sanitization, promotion, demotion |
| City polish replacing agent quality | Track 0 is schema, viewer, and reports. Gate A still blocks live claims |
| Multi-agent too early | One agent that ships beats a planner-coder-reviewer triangle that argues in a sandbox |
The honest bet: harness + discovery + persistence + a spatial report is still underbuilt relative to “wrap the latest model in a chat UI.” Cursor is iterating on the harness in public. OpenAI is standardizing sandbox agents and skills. The open space I want is a cloud loop that remembers, learns discrete policies from PR outcomes, and makes the work visible as a city you can share.
If that city is only a skin, the project failed. If the city is how a run becomes understandable — files, tests, harbor, failures — then the visualization is doing engineering work, not marketing.
That is the plan. The fixture city exists. Next the cloud loop has to earn the harbor: a real sandbox run, a tested change, a PR, and a trajectory that lets us explain what happened.
The next time a cloud agent opens a PR while your laptop is closed, the interesting question is not only whether the patch is right.
It is what the system remembered from last Tuesday, what it refused to load into the window, and whether you can point at a building and see the work.
Sources and further reading14 references
- Dynamic context discoveryJediah Katz, Cursor Blog, 6 Jan 2026
Static vs dynamic context; long tool outputs as files; chat history as files after summarization; Agent Skills; MCP tool descriptions as files; 46.9% fewer total agent tokens on MCP-calling runs; terminals as files.
- Continually improving our agent harnessCursor
Less static preload; agents start light and discover.
- Improving agent with semantic searchStefan Heule, Emily Jia & Naman Jain, Cursor Blog, 6 Nov 2025
+12.5% average question accuracy with semantic search (range 6.5%–23.5%); grep + semantic together best; trace-trained embeddings; code retention and follow-up A/B results.
- Dropbox uses Cursor to index over 550,000 filesCursor
Indexing at scale is a product problem. The model does not eat the whole corpus.
- The next evolution of the Agents SDKOpenAI
Model-native harness, sandbox execution, skills / AGENTS.md / shell / apply-patch, providers including Daytona.
- Sandbox agentsOpenAI Developers
SandboxAgent still has tools, handoffs, and guardrails; the execution boundary is the sandbox.
- Using the OpenAI Agents SDK with Daytona SandboxesDaytona
Shell and filesystem in-sandbox; pause_on_exit / resume; isolated cloud execution.
- Equipping agents for the real world with Agent SkillsAnthropic
Skills as files with a name and description; load full instructions on demand.
- A Contextual-Bandit Approach to Personalized News Article RecommendationLi, Chu, Langford, Schapire, WWW 2010
LinUCB: discrete arms plus context, without training the underlying model.
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?Jimenez et al., ICLR 2024
Opening a PR is not enough. Held-out eval still matters.
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringYang et al., NeurIPS 2024
Interface and harness design change agent success, not only the base model.
- React Three Fiber documentationPoimandres
React renderer for Three.js — the city viewer stack.
- GitHub Apps documentationGitHub
Installation, permissions, PR creation — delivery is GitHub-native.
- Model Context ProtocolAnthropic / MCP
Why tool catalogs explode; background for Cursor's MCP token result.