Video summary

More Speed & Simplicity: Practical Data-Oriented Design in C++ - Vittorio Romeo - CppCon 2025

Main summary

Key takeaways

Educational

Main ideas, concepts, and lessons

1) Data-oriented design (DoD) vs OOP (OP) mindset

  • Speaker’s framing: A common OOP approach models a world of autonomous objects (entities) that own their data and perform their own behavior via virtual methods and polymorphism.
  • DoD alternative mindset: Treat the program as a pipeline of data transformations:
    • Data is the centerpiece.
    • Behavior (update/draw) is centralized in a “world/system” that processes batches of data.
    • Data layout and access patterns are primary drivers of performance and simplicity.

2) Performance comes from memory layout, not just algorithms

  • Key claim: If the data layout causes cache misses, the CPU may spend most time waiting for memory.
  • Cache line concept:
    • The smallest unit moved between memory and CPU cache is a cache line (typically ~64 bytes).
    • Even if you need a few bytes, the CPU still loads an entire cache line.
  • Practical implication:
    • Contiguous “flat” storage improves spatial locality.
    • Predictable access patterns enable hardware prefetching.
    • Sometimes a “worse” algorithm can outperform a “better” one due to memory behavior.

3) The rocket demo as a controlled experiment

  • What the demo simulates: Many rockets with attached smoke/fire emitters and particles.
  • Requirements:
    • Physics-like integration (position/velocity/acceleration).
    • Particles and emitters support time-varying properties (opacity, scale, rotation).
    • Each entity/particle is identifiable via identity/handles.
    • System must be extensible (new actors/effects).
  • What changes in the experiment: not the operations, but the data layout.
  • Observed outcome: an OOP-style layout tanks frame rate; a data-oriented layout makes it playable.

Methodology / implementation path shown (detailed)

Step A: Implement using OOP-style hierarchy (baseline)

  • Create polymorphic base type:
    • Entity with:
      • virtual update(delta_time)
      • virtual draw(render_target)
      • virtual destructor
      • state fields like position/velocity/acceleration
      • reference to World to query other entities / spawn entities
      • an alive boolean to signal removal
  • Derive concrete types:
    • Particle:
      • fields like scale/opacity/rotation and rates of change
      • alive becomes true only while opacity > 0 (fade-out cleanup)
      • overrides draw to render smoke/fire textures
    • Emitter:
      • stores spawn timer/rate
      • virtual spawn_particle() overridden by SmokeEmitter / FireEmitter
    • Rocket:
      • has attached emitters (smoke + fire) and moves with physics
  • World stores:
    • std::vector<std::unique_ptr<Entity>> (heterogeneous heap allocations)
  • Update/draw loop:
    • Iterate all entities in the vector and call update() and draw().
  • Cleanup:
    • Use erase/remove-style logic with predicates (mentions C++20 erase_if).

Why it performs poorly (as explained):

  • Scattered heap allocations → many cache misses.
  • Virtual dispatch overhead (vtable lookups).
  • Frequent dynamic allocation churn when spawning/destroying.

Step B: Move toward DoD by flattening and decoupling (reduce indirection)

  • Goals of the refactor:
    • Remove per-entity heap allocations.
    • Remove inheritance polymorphic trees.
    • Decouple data from logic:
      • Entities become plain structs (data only).
      • The world/system becomes the behavior executor.
  • Replace single heterogeneous container with multiple homogeneous containers:
    • std::vector<Particle>
    • std::vector<Rocket>
    • std::vector<optional<Emitter>> (slot array)
  • Handle emitter/rocket relationships using indices, not pointers:
    • Store emitter references in rockets as indices.
    • Use optional “slots” so indices remain stable:
      • if an emitter is destroyed, clear the slot but don’t reorder.
  • Emitter spawning:
    • World loops over emitters by type (initial version uses an enum-like branching per type).
    • When emitting, push particles into the particle vector.
  • Updates become batch operations:
    • World loops through all particles and applies physics/integration in bulk.
    • World updates emitters and spawns particles from timers.
    • World updates rockets and positions attached emitters using indices.

Benchmarked improvement:

  • Same operations, different data layout → large speedup (about 70% less update time on average in their benchmark set).
  • Rendering also improves (mentioned: bulk sending to GPU).

Step C: Reduce branching and shrink hot data (still AOS initially)

  • Optimization idea:
    • Avoid branching in hot loops by grouping data by type before processing.
  • Implementation:
    • Split particles into separate containers:
      • one for smoke particles
      • one for fire particles
    • Split emitters similarly.
    • Use lambdas/local callbacks to reduce repeated code while staying explicit.
  • Reduce common type sizes / padding:
    • Example: change some fields from size_t-like to uint16_t when safe (to fit more in cache lines / reduce footprint).
  • Result:
    • Smaller improvement overall (~5.7%) in update time (since particles are main bottleneck and branching wasn’t dominant there).
    • Rendering improves further (mentioned: ~12%).

Step D: Change data layout to SOA (structure of arrays)

  • Transform from AoS (array of structs) to SoA (struct of arrays):
    • Instead of Particle { pos, vel, acc, opacity, scale, ... } per element,
    • Use ParticleSOA containing separate vectors per field:
      • positions[], velocities[], accelerations[], etc.
  • Implicit invariant:
    • All field vectors are the same length and grow/shrink in sync.
  • Execution:
    • Update loop iterates by index and updates all relevant vectors.
  • Cleanup in SOA:
    • Standard algorithms don’t fit well, so implement a custom erase:
      • Use a predicate to determine which indices should be removed (opacity fade-out).
      • Perform a partition-like movement:
        • move kept elements from left/right,
        • swap/move across all vectors simultaneously
        • then resize all vectors at once (to keep them aligned)
  • Benchmarked result:
    • Large update-time reduction (~32% decrease).
    • Rendering becomes ~2x faster by mapping grouped fields to GPU buffers more directly.

Why SOA helped even in a “worst case” scenario:

  • Explained overhead from AOS vs SOA vectorization differences:
    • AOS may require gather/scatter when applying SIMD to strided data.
    • SOA tends to enable more continuous SIMD-friendly access patterns.

Step E: Maintain simplicity despite SOA awkwardness (proposed techniques)

  • Problem:
    • SOA makes inserting/updating particles verbose (e.g., multiple push_back calls per field).
  • Suggested direction (not fully implemented in the talk):
    • Create a “magic” wrapper/template that:
      • accepts a “particle-like” aggregate
      • automatically distributes fields into the SOA vectors
    • Use compile-time reflection-like techniques:
      • In principle: C++20/17 techniques (e.g., boost::pfr) for aggregate field access
      • Anticipated benefit from C++26 reflection for more ergonomic APIs using parameter names/types.

Additional lessons: correctness, extensibility, and trade-offs

“SoA is not always the answer”

  • DoD is a mindset:
    • Data layout and access patterns drive the choice.
    • Consider:
      • fields accessed together frequently → store together
      • cold fields accessed rarely → store separately
      • target platform constraints (cache size, SIM length, etc.)
  • Hybrid experimentation:
    • Switching layouts at compile time can allow benchmarking without rewriting logic.

Don’t reject OOP entirely

  • OOP can be valuable at the wrong level vs right level distinction:
    • If OOP is used to wrap data-oriented engines (e.g., “particle manager” API), it can be beneficial.
    • Virtual overhead can be negligible when calls aren’t in hot loops.
  • Team collaboration:
    • Clear interfaces and SRP (single responsibility) can improve understanding and parallel work.

Side benefits of DoD

  • Serialization becomes easier:
    • If state is “just bytes in structures/arrays,” saving/loading and networking snapshots are straightforward.
  • Testability/debugging/tooling:
    • Save and replay state for deterministic repro.
    • Build tools/UI since state is explicit data.
  • Multi-threading:
    • Centralized batch loops can be more easily parallelized than polymorphic object callbacks.

Q&A highlights (key points only)

  • Allocators as intermediate solution:
    • The speaker believes allocators could significantly recover some performance for OOP designs.
  • Loss of explicit relationships (pointer-based ties):
    • Indices make relationships explicit during update; relationships become visible in the central data-processing loops.
  • Testability comparison:
    • DoD helps by making it easy to store/load test cases as data.
    • OOP can still be fine when used for higher-level abstractions; performance-critical storage should remain data-oriented.
  • How to find where to optimize:
    • Use profilers (Intel VTune, Valgrind tools, perf) to detect cache/memory bottlenecks and CPU idle time.
  • Views/ranges overhead and practical use:
    • Ranges/functional abstractions may add overhead without optimization; efficiency depends on inlining/compile flags.
  • Random/non-batch access:
    • Depends on application; can still improve cache friendliness by aligning data layout with access patterns (even for graph-like workflows).

Speakers / sources featured (as identified in the subtitles)

Speaker

  • Vittorio Romeo (main speaker; keynote by Vittorio Romeo)

Referenced people / sources (mentioned as examples or in recommendations)

  • Jason (addressed by name; appears to have introduced/asked a prompt at the start)
  • Mike Acton (gave a keynote in 2014; cited as inspiration)
  • John Leos (co-author)
  • Russell Lapenov (co-author)
  • Alistair Meredith (co-author)
  • Scott Meyers (recommended talk: “CPU Caches and Why You Care” from 2014)
  • Jonathan Mueller (recommended talk at CppCon: “Cache Friendly C++”)
  • Barry (referenced as giving a relevant talk about SOA/reflection; “Barry’s excellent talk on Monday”)

Libraries / standards / tools mentioned

  • SFML (render target; also modernizing to C++17 mentioned)
  • SDL
  • C++ standard / features (e.g., std::erase_if, C++17/20, “reflection in C++26”, std::optional, std::vector)
  • Intel VTune (profiler)
  • Valgrind (suite/profiling tools)
  • perf
  • OpenMP (suggested for parallelization)
  • Open source / tech context: SFML, SDL, Steam, Factorio (for test/state replay example)
  • boost::pfr (mentioned as a way to do aggregate reflection-like behavior in C++17)

Rate this summary

Your feedback will help improve summaries.

Improve this summary

Reprocess with a stronger model when the summary feels incomplete or inaccurate.

Pro

Translate summary in another language

Pro

Ask questions to this video

Chat for follow-up questions, clarifications, and source-backed answers.

Coming soon

Share this summary

Original video