25 min

Astra Solo vs. Astra + GPT-5.6 Sol: Two Three-Hour ISUCON11 Studies

isuconcodexai-agentsperformanceaws

The ISUCON11 qualifier asks competitors to speed up ISUCONDITION, a web service that stores chair-condition reports and serves lists, histories, graphs, and trends. Its benchmark combines continuous writes with reads and aggregation, checks required behavior, and reports a performance score.

I ran two three-hour comparisons between Astra working alone and Astra coordinating GPT-5.6 Sol workers. This article puts the controlled rerun first and retains the initial study below with its original measurements and figures.

Controlled ISUCON11 rerun: Astra solo reached 556,115 during the trial and 439,762 after reboot; the Astra and Sol team reached 409,477 and 429,895

Controlled rerun: 11:14–14:14 JST on September 13, 2026

In the completed 180-minute rerun, I did not observe an efficiency advantage from the coordinated setup. Solo measured more candidates, used fewer recorded tokens, and achieved the higher in-window score. After both selected versions were rebooted and measured once, however, the scores were only 2.3% apart. One independent trial per arm cannot establish that this small gap is significant or repeatable.

Results

MetricAstra soloAstra + Sol
Valid baseline1,6271,843
Best valid score within 180 minutes556,115409,477
Selected version after reboot439,762 / pass429,895 / pass
Normal measurements, pass / fail40 / 323 / 1
Diagnostics, pass / unfinished3 / 17 / 0
Recorded total tokens40,948,90046,719,114
Uncached input1,018,3801,718,097
Output177,176390,713

Solo's in-window peak was 1.358 times the team's. Its reboot result fell 20.9% from that peak, while the team's rose 5.0%; reboot, initialization, host and load variation, and selecting a maximum cannot be separated by one follow-up run. The reboot scores therefore answer a narrower question: both selected versions still passed after a full restart, and their observed scores were close in that measurement.

Controlled rerun score progression, benchmark outcomes, and recorded token use for Astra solo and the Astra and Sol team

Solo reached 10,000, 50,000, and 100,000 first. The team reached 250,000 about 1.1 minutes earlier, but never reached 500,000; solo did so at 158.08 minutes. The 180-minute average of the running best was 238,620 for solo and 200,449 for the team. These are paths from this run, not expected results for future trials.

SDK parent and managed workers

Both parents were Astra/high sessions run through the official Codex Python SDK, openai-codex 0.154.0, using ChatGPT Pro authentication. The coordinated arm's children were managed Sol/high Codex CLI workers. This was not an Agents SDK application calling API models, and the team did not use Codex native subagents.

The parent gave each worker a hypothesis, current baseline commit, owned paths, required evidence, and validation method. Workers ran in dedicated clones with operating-system restrictions on writable paths, networking, and Git metadata. A host-side broker exposed only pinned control and experiment operations. The parent collected each patch, inspected and committed it, obtained independent review, and submitted measurements through a serialized executor. The two-worker cap was enforced. Stale work based on an old baseline became a new attempt, while revisions of the same hypothesis and baseline reused the same thread.

The team used 29 child threads in 34 invocations, including five resumes. Children overlapped for only 8.00 minutes. A completed read-only review with no diff remained in prepared state and occupied one slot for 28.01 minutes; two launches failed because the cap appeared full. This proves lost worker capacity, but the experiment did not measure a counterfactual score loss. The control software was kept fixed and was not repaired during the trial.

Managed Sol worker intervals and serialized benchmark intervals during the controlled 180-minute rerun

Score, tokens, and cost over the run

The following figure relates the running best valid score to elapsed time, cumulative recorded tokens, and allocated EC2 instance-hours. A flat score while tokens increase means that expenditure has not yet produced a new measured best; it does not mean the investigation was worthless.

Rerun score against time, cumulative recorded tokens, and allocated EC2 instance-hours; monetary cost is not measured

Usage updates are merged chronologically across threads, without counting resumed threads twice. The curves use reported event times rather than assuming constant token consumption between events. Only ordinary passing runs completed within the deadline update the best score. Missing usage from the interrupted final turns is not imputed. Open the full-size PNG or PDF, or download the figure CSV.

Dollar cost was not measured per arm. I did not allocate an arbitrary fraction of a Pro monthly subscription to these three hours, or multiply Codex usage by API prices and label it a bill.

Resource or expenseRecorded comparisonUnavailable measurement
LLM useParent/child and model-level token usagePer-arm Pro allowance consumption or incremental charge
AWS allocationFour instances per arm for three hours: 12 instance-hoursInvoice-reconciled dollars
Setup and final verificationOutside the 180-minute optimization windowTotal per-arm expense including these stages

The EC2 axis is a resource-time proxy: nine application-instance-hours plus three benchmark-instance-hours per arm. Because both arms have the same machine mix, it rescales the time axis and supplies no independent evidence of a monetary advantage. A dollar curve would need applicable machine rates, actual lifetimes, storage, and network charges. Cached input, uncached input, and output also cannot be treated as equally priced units.

Why did using the Codex SDK not produce a better result?

The official Codex SDK documentation, checked on September 13, describes the Python SDK as a JSON-RPC client for the local Codex app-server. The ISUCON task ledger and acceptance policy in this experiment were our own implementation.

This study found that our coordinated setup did not beat solo; it did not isolate a negative SDK effect. Both rerun parents used the SDK. The earlier comparison also changed too many controls together to answer the SDK-only question.

Evidence or policyResponsible layerSupported interpretation
A completed review held a slot for 28 minutes; two launches were blockedOur coordinator lifecycleProven capacity loss, not an SDK failure; score impact unknown
Two workers overlapped for only eight minutesTask decomposition, dependencies, and dispatchLimited realized parallelism; unoccupied time is not automatically waste
The parent used 87.9% of recorded tokensCentralized decisions, integration, and reviewSmaller worker context did not remove parent work; latency attribution is unmeasured
Team produced 23 ordinary passing measurements versus 40 soloThe complete implementation-to-measurement workflowFewer validated experiments, without identifying the cause of the score gap
Workers passed network-isolation checks; no logged direct parent implementation; pending candidates were boundedAdded policies and retained audit evidenceEach restriction's performance effect is unmeasured; unlogged effects cannot be ruled out

The SDK let us control persistent threads, collect events and usage, and interrupt work at the deadline. Switching interfaces did not by itself improve hypothesis selection, independence between tasks, or review granularity. BrokenPipeError and TransportClosedError at cutoff were recorded during deadline shutdown; they are not evidence of an earlier failure explaining the whole trial.

The next proposed change is to terminate zero-diff review tasks correctly and count running workers separately from pending candidates, validated before another timed run. To isolate SDK overhead, keep the CLI binary, model, prompt, permissions, tools, and environment fixed, changing only direct CLI versus SDK delivery; measure first-response latency, tool round trips, and completion time. Scheduling changes belong in a separate comparison. These are proposed tests, not improvements already demonstrated here.

Conditions, accounting, and limits

Each arm used three c5.large application instances and one c4.xlarge benchmark instance, for eight independent EC2 instances in one availability zone. They came from the same reconstructed AMI rather than the unavailable bit-identical competition image. Seed source and deployed binaries were matched. Four preparation runs against the AMI's bundled binary were marked invalid and excluded; after rebuilding the seed, each arm received one valid baseline run. Solo's baseline passed with one nonfatal deduction.

The team recorded 46.72 million total tokens versus solo's 40.95 million. Its Astra parent accounted for 87.9% of that recorded total and reached a 244,705-token request, compared with the largest child's 53,083. Both root sessions were interrupted during their final turn at the deadline, so the totals are observed lower bounds and may omit that unfinished work; they are not billing records.

All eight machines changed boot ID during final verification, and all six running application hashes matched their selected builds. Both reboot benchmarks passed with zero deductions. Destruction was then verified for the eight instances and associated resources.

The sanitized rerun summary and score CSV are available for inspection. The earlier study's best in-window scores remain 1,621,980 for solo and 1,114,247 for the team. Because the rerun changed the SDK controller, broker, isolation, concurrency, ledger, parent role, and context handoff together—and hardware and search paths also varied—the lower absolute scores do not identify an SDK effect or support a general model ranking.


Initial study: 00:09–03:09 JST on September 13, 2026

Does adding coding agents make performance optimization faster? I gave one Astra agent and an Astra coordinator with GPT-5.6 Sol children the same starting code and three hours each in Codex.

After rebooting the servers, the solo run scored 1,636,308, compared with 885,334 for the team: about 1.85 times higher. Yet the team reached 500,000 points roughly 27 minutes earlier. Neither the final score nor a favorable midpoint tells the whole story.

ISUCON11 qualifier, one 180-minute trial per arm. Final scores: solo 1,636,308 and team 885,334. Total tokens: 61.47M and 82.33M

This article is based on the experiment report, experiments/codex-aws-compare/RESULT.md. The measured task was the ISUCON11 qualifier. Below, I use the original report figures and recorded execution data to examine scores, progress, tokens, and coordination.

What is ISUCON?

ISUCON is a competition in which participants improve an existing web service on limited server resources while preserving its required behavior. Changes can include reducing database queries, making computation cheaper, caching results, or redistributing work between servers.

A benchmark program simulates users, checks behavior, and assigns a performance score. The points in this article are its output, not raw requests per second. The scoring rules depend on the problem.

The measured problem was ISUCONDITION, from the ISUCON11 qualifier. It receives chair condition data and displays histories, graphs, and trends. Continuous writes compete with reads and aggregation, making databases, CPU work, caches, and server placement useful optimization targets. See the official application manual.

I wanted to compare the ability to deploy working changes and produce measured improvements, beyond producing convincing explanations.

Eight EC2 instances, two independent environments

Each arm received three competition servers and one benchmark server. The two environments ran simultaneously, without sharing application databases or benchmark machines.

ConditionSoloCoordinated team
Parent modelAstra / highAstra / high
ChildrenNoneGPT-5.6 Sol / high
Optimization time180 minutes180 minutes
Independent trials11
Competition servers3 × c5.large3 × c5.large
Benchmark server1 × c4.xlarge1 × c4.xlarge
Starting codeSame commitSame commit

The model identifiers in the logs were gpt-6-astra and gpt-5.6-sol. Both used high reasoning effort and Codex CLI 0.154.0. The window was September 13, 2026, 00:09:35–03:09:35 JST. Setup and baseline measurement happened beforehand. This was an equal-time comparison, not an equal-token-budget comparison.

I matched the official ISUCON11 machine configuration: three c5.large competition instances and one c4.xlarge benchmark instance per arm. All eight ran in the same availability zone, each with 20-GB gp3 storage at 3,000 IOPS and 125 MiB/s. Benchmark machines were not upgraded during the trial.

The original official AMIs were unavailable. I used matsuu's reconstructed environment, matching the instance types, counts, and 20-GB gp3 storage. This was not a bit-identical recreation of the original competition image. Physical CPU models also differed despite matching instance types. Baseline scores were 1,678 and 1,874.

The shared starting commit was 0652d01d2699135605d8b634531d9873e9d71b2e. Application, SQL, public-file, certificate, and benchmark-executable hashes were checked across the environments. Neither arm received a previous optimized solution. Shared inputs and execution controls were frozen before the trial, and no optimization advice or results from the other arm were introduced during it.

Each CLI finished once about two minutes before the deadline. I resumed the same thread once and stopped it at the fixed deadline; this did not add an independent trial.

How the team was organized

The team received an explicit operating policy:

  • Start with separate investigations of the database, APIs and requirements, and runtime behavior.
  • Split implementation by testable hypothesis. Allow at most two implementation tasks and five children overall at once.
  • Give implementation workers separate Git worktrees and focused starting context.
  • Return summaries and files for the parent to integrate.
  • Serialize deployment, initialization, benchmarking, and log collection within each environment.

The five-child limit was explicitly set by the AWS trial controller. It differed from the repository’s ordinary .codex/config.toml, which still contained a two-child limit from an earlier small comparison. The CLI ignored user configuration; agents were disabled for solo. The audit checked actual launch settings and native sessions, rather than inferring behavior from a configuration file alone.

A Git worktree provides a separate working directory for code changes. It does not isolate a database or a running service, so code isolation and environment locking were separate mechanisms.

The executor used the existing isucon-harness make aws-deploy and make aws-bench-only commands. It pinned each candidate commit and checked the running executable hashes on all three application hosts before measuring. Decisions used public manuals, benchmark output, and our own logs. Benchmark source code was not read.

How parent–child coordination actually works

There are three control layers. A Python controller owns trial timing and CLI settings, the Astra parent assigns work, and a Python executor serializes measurements on the servers. The controller does not decide which SQL query to optimize next; hypothesis selection and delegation belong to the parent.

ComponentResponsibilityOutput
controller.pyLaunch both CLIs with fixed timing, models, common task, and team policyRoot threads, JSONL events, exit/resume records
Astra parentChoose hypotheses, spawn and message children, review and integrate changesBounded child tasks and candidate commits
Sol childrenInvestigate, implement in dedicated worktrees, or review a diffConclusions, evidence files, commits, validated and unverified points
executor.pyAccept a commit from the parent; deploy and measure under an environment lockRun ID, score, messages, manifest, logs
session_audit.pyFollow recorded parent–child relationships, models, tool calls, and usageActual child counts, overlapping turns, token totals

1. Fix models and concurrency at launch

The following TOML excerpt summarizes the relevant settings passed by the controller. These were used with CLI 0.154.0 in this experiment; the excerpt alone is not a complete experiment setup.

model = "gpt-6-astra"
model_reasoning_effort = "high"
 
[agents]
enabled = true
max_concurrent_threads_per_session = 5
default_subagent_model = "gpt-5.6-sol"
default_subagent_reasoning_effort = "high"

The parent uses Codex’s collaboration.spawn_agent to start children and collaboration.send_message to send follow-up information. Children are instructed not to launch separate Codex CLIs, keeping work within a traceable thread tree. The record contains 24 spawn calls, 22 sends, and no rejected spawns.

2. Send a task contract, not the entire conversation

multi-policy.md requests fork_turns="none" for every child. Each receives an objective, current baseline commit, hypothesis, owned paths, observation references, required artifact, validation method, and prohibition on external mutations. This illustrates that handoff format; it is not a verbatim historical prompt.

Objective: Test whether graph retrieval can read less data from the DB
Baseline: The current commit specified by the parent
Scope: Graph retrieval only; preserve writes, authorization, and scoring rules
Evidence: SQL logs from this trial's run and the relevant handler
Work: Create a dedicated worktree at the baseline; fetch only required columns
Return: Commit, evidence, local validation, unverified points, and risks
Do not: Deploy, change remote DBs, restart services, or run benchmarks directly

Children return short responses and files, and the parent reads the relevant evidence. Commits, diffs, observations, and task notes carry shared state; there is no continuous synchronization of everyone’s conversations into a shared memory. Sharding work actually used separate design, implementation, and fresh-review stages. The implementation child owned application condition routing, while the parent prepared configuration and initialization for the two database hosts. The parent integrated both into one candidate, and a separate child reviewed SQL, consistency, and caching before measurement.

3. Integrate centrally and overlap independent work with measurement

The policy allows at most two independent implementation tasks and one integrated but unmeasured candidate. When a child returns a commit, the parent reviews its diff against the hypothesis and API contracts, then integrates it onto the current baseline. A spare child slot can be used for a focused review.

When the baseline changes, the policy calls for notifying only children whose paths or assumptions are affected, sending the new commit and changed conditions. Dependent candidates must be brought onto the current baseline and measured after integration. Other independent investigation, implementation, or review can continue during a benchmark. As the timeline shows, this overlap was underused in the actual trial.

4. Route server operations through one entry point

The parent submits a candidate with ./experiment.py candidate --commit COMMIT_SHA --label SHORT_DESCRIPTION. Read-only runtime observations use the same client’s inspect operation. Children return candidates; the parent is the sole evaluator operator under the policy.

For each environment, the executor holds flock across the candidate snapshot, baseline configuration reset, candidate operations, build/deploy, running-binary verification, benchmark, and evidence collection. It also checks for a previous benchmark still running. The lock belongs to this execution path, rather than to an LLM conversation that may pause or resume.

Scores and pass status are parsed from execution output. Execution errors or missing required evidence produce an incomplete record. An interrupted benchmark whose remote cleanup cannot be established can quarantine the environment, blocking the next experiment. The parent’s final prose is not the source of the official score.

The code controls CLI settings, deadlines, and the locking and records of operations routed through the executor. The semantic two-implementation limit, child write ownership, and prohibition on direct SSH depend on instructions and retrospective audit. They are not all enforced by per-child operating-system permissions. That distinction bounds what “controlled” means in this experiment.

The team led in the middle; solo overtook it late

Original report: best eligible score, score divided by each arm's baseline, cumulative total tokens, and uncached input/output. Blue is solo; orange is the team

Full-size PNG / PDF. This is the original report figure, unchanged. Blue denotes solo; orange denotes the team.

The top left shows the best normal passing score observed by each time; the top right divides that score by each arm’s own baseline. Both exclude diagnostic profiling runs, failures, and the post-window reboot benchmark. The bottom left shows cumulative tokens including cached input; the bottom right separates uncached input and output.

ThresholdSoloTeam
10,0000:02:430:04:51
50,0001:00:020:58:14
100,0001:06:471:16:09
250,0002:17:101:58:54
500,0002:31:052:04:05
1,000,0002:38:232:37:29

The team reached 500,000 about 27 minutes sooner. At one million, its lead had narrowed to 54 seconds. Solo then made larger gains.

Integrating the running best passing score over time and dividing by 180 minutes gives 316,651 for solo and 370,073 for the team: a 16.9% team advantage. This summarizes how early high scores were observed. It does not establish the service’s actual performance at every instant or the repeatability of those scores.

The selected solutions differed, too.

AreaSolo selectionTeam selection
WritesGrouped transactional LOAD DATA; respond after commitINSERT all conditions in transactions of 100 records
Placementapp3 hosts the DB; app2 handles ingestion; app1 owns metadata/trendDeterministic condition sharding across app2/app3; metadata on app2; apps on all three
ReadsCondition index, fewer columns, 16-second trend cache, reused Gob payloadsCondition index, graph projection, 250-ms trend cache, icon LRU/ETag
Selected commit, shortened0e2a05019f669e7dae4eee86

This compared optimization paths as well as agent counts, rather than the time to produce an identical patch. Solo gained late through trend TTL and write placement; the team gained in the middle through database sharding and graph projection.

Solo's final trend cache lifetime was 16 seconds; the team's was 250 milliseconds. The public requirements permit delayed condition updates in the trend view. This is an observed difference in implementation, but I did not isolate cache lifetime as the sole cause of the outcome.

Peak scores and final reboot scores are different measurements

MeasurementSoloTeam
Best normal score within three hours1,621,9801,114,247
Pre-deadline confirmation of the selected commit1,608,134654,392
Final score after reboot1,636,308885,334

The first baseline loads started at 23:48:04 JST for solo and 23:48:01 for the team on the previous day. Solo reached its timed peak 3:12:23 after that first load, or 2:50:52 into optimization; the team took 2:59:03 and 2:37:29 respectively. Final reboot runs completed at 3:23:42 and 3:23:46 from the first load, outside the optimization window. No idle intervals were compressed.

“One trial per arm” means one three-hour optimization attempt. Each attempt contained successive candidate measurements, including confirmation of the selected version.

Following the ISUCON11 manual, I used the post-reboot benchmark as the final score. Both passed, although the team had nine transport errors and 71 timeouts. Its output was 885334(885350 - 16): the direct penalty was only 16 points, so penalties alone do not explain the score gap.

All four hosts per arm changed boot IDs, and the three application executable hashes were checked again after reboot. Solo had zero penalties and zero timeouts, but its stderr included a Force ending loadWaitGroup warning, which was retained.

The team varied substantially on the same selected commit. Rebooting cannot be identified as the only cause of the drop. A further persistence check—rebooting again and reading back data written during the final benchmark—and a browser equivalence replay were not performed. The runtime evidence establishes a passing benchmark after reboot, not completion of every official follow-up check.

Count the children's tokens, too

UsageSoloTeam parentSol childrenTeam total
Input61,273,76054,406,27227,412,20481,818,476
Cached input60,027,39253,330,04825,752,32079,082,368
Uncached input1,246,3681,076,2241,659,8842,736,108
Output193,901208,014308,477516,491
Reasoning output96,503103,446108,485211,931
Total61,467,66154,614,28627,720,68182,334,967

Cached input is a subset of input. Total is input plus output; cached tokens are not added again. Reasoning output is likewise counted within output.

The team's parent used less input than solo. Including the children, however, total tokens increased by 34%, uncached input by about 2.20 times, and output by about 2.66 times. Reducing the parent's context burden did not reduce the system's total usage.

The team parent used 11.2% less input, but recorded compactions were two for solo and three for the team parent. Peak request input was 244,601 tokens for solo and 218,111 for the team parent; one research child reached 193,828. These are recorded request sizes, not model context-window capacities.

Fresh context did not guarantee small requests. All 24 children had references to reading the manual, so repeated onboarding remained part of the workload.

Usage came from all 26 experimental threads—one solo thread and 25 team threads—in native session logs. Preparation, observer, and post-experiment code-review agents were outside these totals. The deadline interrupted the parents' resumed turns, leaving incomplete turn records. These totals reflect recorded usage, not a finalized bill covering any unreported provider-side processing. I did not assume model prices to convert them into dollars.

Twenty-four children did not mean continuous parallel work

Original report: 25 completed turn intervals across 24 Sol children, with the serialized benchmark lane below and the 180-minute deadline dotted

Full-size PNG / PDF. Orange marks child turn intervals; the bottom blue lane marks benchmark runs. The three initial investigations overlap, but much of the later activity uses one child at a time.

The logs established that:

  • Solo launched no children. The team launched 24 child threads, all GPT-5.6 Sol with high reasoning effort.
  • Peak concurrent child activity was three. No grandchildren were launched.
  • All children used fork_turns="none". The recorded commands showed separate implementation worktrees, with no counterexample to the two-implementation limit.
  • There were 25 completed child turn intervals. No child was active for 76:49, one for 88:39, and two or more for 14:32.
  • Summed child activity was 119:37, with an average active child count of 0.665.

“Active” here means a child turn interval, including tools and waits. It is not measured GPU inference time.

Model selection, counts, and execution intervals were auditable. Worktree ownership and meaningful task independence were not enforced as operating-system security boundaries. Observed compliance is narrower than mechanical enforcement.

Experimental throughputSoloTeam
Candidate executions5656
Normal passes5046
Diagnostics48
Failures22
Normal passes per hour16.6715.33
Numerical best-score updates2411
Executor occupancy87.25 min85.54 min

Total lock waiting was approximately 0.003 seconds per arm. The record does not indicate a continuously saturated executor. Candidate generation, investigation, and integration could be coordinated better, although the remaining time cannot all be classified as wasted waiting. A numerical best-score update is also not a statistically established improvement.

Failures were retained: bounded-history-reads and grouped-durable-writes for solo; trend-single-query and three-durable-homes for the team. Looking only at successful candidates would hide part of the work.

The team parent spent about 17.4 minutes in sleep calls, including waits for children and benchmarks. It also implemented a late pool change itself, rather than acting exclusively as a coordinator. These observations describe the workflow; none independently proves the cause of the final score gap.

What I would change next

For this workload, the result supports keeping Astra solo as a baseline and using children selectively for independent investigation and review. The team's earlier progress is still useful evidence.

I would make the pipeline more explicit: while the parent deploys and measures, children prepare one or two independent hypotheses for the next experiment. A concise shared map of the requirements should reduce repeated reading, while preserving links to the relevant original text. Near the deadline, I would reserve time to examine errors and repeatability rather than continue adding structural changes.

Those are proposed improvements, not a measured winning configuration. This experiment tested one team setup with an efficiency policy, not an established optimum for multi-agent orchestration.

The fixed three-hour deadline also means this was not a run to the harness’s usual optimization plateau.

With one trial per arm, hardware differences, unequal baselines, and different optimization paths, the result does not establish a universal model ranking or a multiplier transferable to other workloads. The operating metric I want to improve is how many useful changes we can measure correctly within a fixed time, rather than the cumulative number of agents launched.

Data and cleanup

The summary JSON and candidate score CSV accompany this article. The opening summary figure was generated from recorded measurements. The four-panel chart and child timeline are unchanged PNG/PDF copies from RESULT.md; their SHA-256 manifest is included. The CSV includes diagnostics, failures, and post-reboot runs; its eligible column identifies normal passing runs within the time window.

I retained 70,450 files, about 7.50 GB of logs, commits, patches, configuration, and parent/child sessions locally, with a per-file SHA-256 inventory. The public data excludes AWS connection details and raw sessions. AWS API checks confirmed that all eight EC2 instances were terminated, all eight associated EBS volumes were gone, and no dedicated security groups or key pairs remained. Destruction and cleanup verification completed at 03:14:14 JST on September 13, 2026.

日本語版