Back to projects February 2026 - July 2026

Counter-Strike 2 Dream Team: Cloning the Pros to Search for the Best Lineup

A machine learning system that trains player-specific behavioral models from professional match demos, tests whether they stay distinguishable on unseen maps, and simulates 37,032 candidate lineups on a GPU to find which five perform best together.

  • Transformers
  • Behavioral Cloning
  • GPU Simulation
  • Sports Analytics
  • Python
Counter-Strike 2 Dream Team: Cloning the Pros to Search for the Best Lineup

If you could sign any five professional Counter-Strike 2 players, which lineup performs best together rather than which five carry the highest individual ratings? I built a system to answer that inside a simulator: it learns player-specific behavior from real match demos, checks that the learned behavior stays distinguishable on maps it has never seen, drops those models into a GPU-batched match engine, and searches tens of thousands of role-valid lineups.

I designed the research question, the validation criteria, and the standard for what counted as a real result rather than a flattering one, then used Claude Code to implement against those constraints. The code is open at github.com/hpagni/cs2-clone-lab.

The result

After screening 37,032 lineups and simulating the finalists across more than a million rounds, one roster ranked first:

sh1ro on the AWP, HUASOPEEK calling, luchov entrying, donk on the flex, and afro in support. It won 76.9% of its simulated matches against a field of the strongest teams in the corpus (95% confidence interval 74.7% to 79.0%).

Disabling the modeled chemistry term and rerunning the same five on individual skill alone dropped that win rate to 64.1%, a 12.8-percentage-point difference. Within the simulator, modeled lineup fit contributed materially beyond individual skill.

Simulated win rate with the chemistry term enabled versus disabled, a 12.8 point difference

The same five players at the same modeled skill, scored with and without the chemistry term. The gap is what the model attributes to fit.

1. Collecting the data

The raw material is professional GOTV demos, the replay files that record every tick of a match, sourced from HLTV and, for deepening individual histories, FACEIT. HLTV serves them behind Cloudflare on short-lived session-bound links, so ingestion runs through a rotating proxy pool with per-IP stickiness to keep the link-resolve and download steps on one network identity.

Each demo is parsed with demoparser2 into two grains. A 4 Hz whole-map layer captures every player’s position, health, economy, grenades, and the kill and round events. A second pass stores a 16 Hz view of all ten players plus a 64 Hz close-up in a two-second window around every shot, carrying both raw view angles and the recoil-subtracted intended aim recovered by undoing each weapon’s recoil pattern. Crosshair placement, flick-versus-track, and reaction time live in that channel, and they are the strongest available fingerprint of who a player is.

The corpus is about 2,045 professional matches (roughly 4,800 map files) plus several thousand FACEIT demos. Per-player depth is thin and uneven, with a median of nine maps per player, and that scarcity shapes the whole model design below. Features are stored as compressed Parquet on Backblaze B2 and queried in place with DuckDB under point-in-time rules, so no model trains on information from after the map it is predicting.

2. Modeling a player

The governing constraint is pure behavioral cloning and never reward maximization. The model only predicts the player’s actual next action, including their mistakes. Adding a win reward, self-play, or sharpening toward the most likely move stops modeling a human and starts building a generic superhuman bot. Skill is treated as a calibration dial matched to each player’s own consistency, never as something to push higher.

Because a player may have only a dozen maps of history, no per-player model could train on its own. The architecture is one large model trained on everyone with a small amount of per-player adaptation on top. A transformer takes all ten players as tokens each frame and attends across time and players at once, predicting a distribution over the next map zones plus a coarse intent such as hold, rotate, or take utility. Predicting a distribution rather than an average matters: averaging “peek left” and “peek right” produces “walk into the wall.” Each player carries a learned identity vector that modulates every layer, so tendencies color the whole network rather than being appended at the end. Specializing a clone means freezing the shared backbone and training only that identity vector and a small low-rank adapter.

Aim is handled by a separate autoregressive network, since crosshair movement is finer-grained than the macro model. It predicts the next movement as a classification over a grid of yaw and pitch steps rather than a single number, because the average of a symmetric aiming error is a perfectly centered shot, which is exactly the superhuman artifact to avoid. Guardrails reject any clone whose aim comes out faster or steadier than the human ever was.

3. Validating player-specific behavior

The validation tests whether a clone’s behavior stays distinguishable as that specific player on maps it has never seen. It does not establish identity in any stronger sense.

Every clone is scored by two evaluators from different model families: a gradient-boosted tree reading hand-built aim statistics, and a convolutional network reading the raw aim sequence. A clone must satisfy both. The criteria ask whether the behavior is recognizably this player rather than a swapped identity, whether it is as distinctive as the real player without becoming a caricature, whether that distinctiveness survives a permutation test with multiple-comparison correction, and whether the aim falls inside the human range.

Two design choices keep the test meaningful. Evaluator weights and acceptance thresholds are fixed independently of clone training, and a sealed block of each player’s maps is held back from every training and evaluator-fitting step, so a clone must pass on both a fresh reserve and that never-consulted set. Several clones passed the reserve and then failed the sealed set. They were recorded as failures rather than re-thresholded.

4. Simulating a match

fastsim runs thousands of full matches at once on a GPU, which is what makes searching a large lineup space affordable. Each round steps through a buy phase, a strategy call, movement, contact, and the duels and bomb plays that decide it, then settles the economy and swaps sides, vectorized across the batch.

The most important rule is fog of war. Every agent acts on a belief about enemy positions built from the last time it actually saw them and how stale that sighting is, never on true positions. Without it the clones would play like wallhackers and nothing would transfer. Movement runs over a per-map navigation graph, duels resolve through sampled reaction time and aim error rather than a fixed coin flip, and a coaching layer keeps the five coordinated.

Running millions of rounds also produces a map of where defensive rounds are won. The first heatmap shows where CT players tend to be; the second shows which of those positions actually hold rounds, with thinly-sampled cells greyed out.

Heatmap of CT player occupancy across simulated rounds on Mirage

Where simulated CT players spend time on Mirage, aggregated over the full simulation run.

Heatmap of CT round win rate by holding position on Mirage

Round win rate by holding position. Occupancy and effectiveness do not coincide: several heavily-used positions hold rounds at below-average rates.

5. Searching the lineup space

Simulating every possible lineup at full fidelity is far too expensive, so the search runs as a funnel. A fast surrogate score ranks 37,032 role-valid lineups, the top 64 go to a first simulation round, the best 16 to a deeper one, and the final 6 are simulated hardest across six maps. Cheap ranking allocates expensive compute to the candidates most likely to matter.

Search funnel narrowing 37,032 lineups to 64, then 16, then 6, then one champion

Each stage simulates fewer lineups in more depth, so compute is spent where it changes the answer.

The surrogate is a transparent chemistry score rather than a learned black box. It rates a lineup on six readable terms: role fit, whether there is exactly one AWP and one caller, playstyle complementarity, prior history together, and map-pool overlap. This is a modeled construct built from those inputs, not a measured property of real teams.

Pairwise chemistry scores among the champion five

Modeled pairwise fit within the top-ranked lineup, by the six-term surrogate.

6. Where the lineup is strong and weak

The champion does not win evenly. It is strongest on Anubis and Nuke, works hardest on Dust2, and still clears 63% against the toughest opponent in the field.

Simulated win rate for the champion lineup broken down by map

Per-map simulated win rate for the top-ranked lineup.

Simulated win rate for the champion lineup against each opposing team in the field

Win rate against each team in the simulated field. The spread is the margin the ranking rests on.

Every figure traces to simulation output, and the confidence intervals come from those same rounds. The full findings, including the positioning heatmaps, best set plays per map, and the chemistry breakdown, are packaged as a self-contained deck in the repository.