Section 2 of the Yellow Paper describes Ethereum as a state machine driven by transactions. The notational conventions in that formalism are: σ for the world state, T for a transaction, Υ for the transaction-level state transition function that applies a transaction to the state; B for a block, Π for the block-level state transition function, and the transactions inside a block written in order as T₀, T₁…. On those conventions the Yellow Paper gives two equations: the state transition of a single transaction is σ_{t+1} ≡ Υ(σ_t, T), and the block-level transition is Π(σ, B) ≡ Υ(Υ(σ, T_0), T_1)…. The second equation deserves a longer look: it writes the execution of an entire block as a nested call of functions, where the input to the second transaction is the state after the first has executed.
Evaluation order is therefore part of the specification, and clients have no latitude to choose otherwise. The questions to pursue are: why it can only be defined this way, what would be lost if two transactions were allowed to be evaluated at the same time, and how this choice locks EVM throughput to a single thread. This piece deals with that layer alone, the design origin.
Order is written into the specification, not left to client scheduling freedom
The Yellow Paper's formulation pins down what the execution layer has to do: given a parent state and a sequence of transactions, apply the state transition function repeatedly to obtain a state. One of the block header's validity conditions is that the root obtained by folding this state through the trie must equal the stateRoot in the header. One transaction finishes executing and its state is fully written back before the next begins reading state; that ordering relation is part of the definition itself, with no room left for optimization.
The execution layer has no ordering authority either. The transaction order is given by the transaction list produced by the consensus layer, and the execution layer simply evaluates in the given order. This division of labor means the consensus layer does not have to agree with the execution layer on “the result after interleaved execution,” only on “which transactions were included, and in what order.” Any node that replays the same block in a different order gets a root that does not match, and its result is ruled incorrect outright.
# Block-level state transition: the input state of the second transaction is the output state of the first
state = parent_state
for tx in block.transactions:
state = apply_transaction(state, tx) # Υ(σ, T)
state = apply_withdrawals(state, block.withdrawals) # Withdrawals execute after all transactions
assert trie_root(state) == block.header.state_root
Hard ordering constraints: nonce, balance, and cumulative gas
Before any contract code is interpreted, the protocol layer already demands order. Among the initial transaction validity checks listed in Section 6 of the Yellow Paper is that a transaction's nonce must equal the sender account's current nonce. Two transactions from the same sender therefore have a protocol-enforced total order: move the later one earlier and the transaction is simply invalid, not a different result. Another check requires the sender account to have no deployed code (EIP-3607), which likewise reads state as it stands after a transaction has executed.
Balance and gas lock in order at the same layer. Every transaction has its sender's balance checked before execution as sufficient to cover the upfront cost, and that balance is the result of the previous transaction's execution; a block's gas limit is a block-level constraint, and the cumulative gas used in a receipt accumulates from the preceding transactions (in its section on block receipts, the Yellow Paper writes the cumulative value of the nth transaction as the sum of the previous cumulative value and this transaction's own usage). How much gas a later transaction can use depends on how much has already been spent.
In other words, even setting contract storage aside entirely, the premise that state has an order is already established in the protocol rules. Contract execution merely keeps stacking dependencies on top of a sequence that is already strung together.
Where conflicts come from: intersecting read-write sets, knowable only after the fact
The standard way to decide whether two transactions can execute at the same time is to compare their read-write sets: the set of state locations a transaction reads or writes. As soon as a location one of them writes is read or written by the other, execution order affects the result, and transactions of this kind are called conflicting.
Conflicts are common under real workloads. A swap transaction on an automated market maker (AMM) pool rewrites the reserves and the price accumulator, and a lending liquidation right behind it needs to read that same pool's price; the order of the two transactions directly determines whether the liquidation triggers and who absorbs the loss. Dependencies like this need no one to manufacture them — any time contracts share state, they appear.
More troublesome still, the executor cannot see the read-write set. The EVM lets a transaction call any address and read or write any storage slot, and both the call target and the slot number can be computed only at runtime:
// Call target and arguments are determined only at runtime; static analysis cannot produce the read-write set
function dispatch(bytes32 poolId, bytes calldata payload) external {
address pool = pools[poolId]; // The target comes from storage; it may be any registered contract
(bool ok, ) = pool.call(payload); // Which slots are touched depends on pool and payload
require(ok, "call failed");
}
A mapping's slot number is obtained by concatenating the key with the slot position and taking keccak256, the key can be a runtime input, and the slot position is not necessarily fixed at compile time either. So “which slots this transaction will touch” cannot be read out of the bytecode before execution; it can only be observed by actually running the transaction. EIP-7928 states this bluntly in its motivation section: without knowing in advance which addresses and storage slots will be accessed, transaction execution cannot be parallelized.
What single-threaded execution buys: determinism, and verification as replay
The payoff of this design is concrete. Every node evaluates in the same order and arrives at the same state root; a verifier does not need to trust a scheduler or reason about the various possible interleavings, only to replay the same block under the same rules and then compare the stateRoot in the header. What lets multiple independent implementations, in different languages and on different hardware, agree on a single value is precisely the compression of uncertainty into “inputs plus order.”
Failure semantics depend on order too. REVERT and exceptional rollback undo layer by layer using the state snapshots recorded during execution, and the granularity of a rollback is the call stack and the transaction; gas refunds, cumulativeGasUsed, and the status bit in a receipt all have a unique meaning only when transaction order is fixed. If several transactions were allowed to advance interleaved, “the part executed first rolls back while the part executed later stays” would need a whole new semantics to define it.
So one-at-a-time serial execution is indeed a trade-off, but what it buys — determinism verifiable network-wide, and a mode of verification that needs no separate correctness proof — is what a public chain is founded on, and explaining it away as “lazy implementation” does not hold up.
Clients do use all their cores — but always outside the semantics
On the engineering side, mainstream clients do not waste the machine. Take Reth: transaction signature recovery is handed to a thread pool for parallel execution; hashing state updates and constructing the state root are handed to parallel sparse-trie tasks, with multiple workers sharing proof generation and trie node reads and falling back to serial computation when a task times out or fails; and along the block processing path transactions are also executed ahead of time in a separate thread pool to prewarm caches for the later canonical execution.
The canonical execution itself is still sequential. In Reth's block processing flow, the default path without a block-level access list performs sequential EVM execution in block order and hands state updates to the state root task as a data stream; a block carrying a BAL has a separate parallel execution branch, but it depends on the Amsterdam fork and on the presence of a BAL (the block-level access list path discussed below). Prewarm execution only fills caches, and its results are never committed directly. That line was drawn by the design origin: multiple cores can speed up signature verification, hashing, proof generation, and cache filling, but they cannot shorten “the one evaluation chain that is semantically unique.”
One phrase is worth clearing up along the way: single-threaded does not mean only one core of the machine is working; it means there is only one evaluation order in the semantics. That is also why the key question for scaling is not the number of cores but whether “this evaluation chain can be made wider.”
The cost: parallelism ceded outside the protocol
The specification fixes the order, so if the execution layer wants to advance transactions on multiple cores, the only legitimate direction is to find an interleaving equivalent to the canonical order and converge the result back to the same state. That requires knowing or discovering the read-write set first: either the transaction's originator or the block builder declares it in advance, or the runtime tracks it dynamically and handles conflicts as they arise. The former pushes the burden onto the protocol and tooling; the latter pushes it onto the runtime and rollbacks.
The premise of global state makes the cost more visible. The state is a single tree shared by all nodes, every transaction's update rewrites nodes along the path from leaf to root, and state access itself becomes a link on the critical path. Disk I/O and network bandwidth can scale out; the chain of “fetch state, compute the result, write state back, then fetch the next transaction's state” cannot.
The direct answer: declare the read-write set in advance
The direct response to this origin is to supply the missing information. The block-level access list (BAL) proposed in EIP-7928 adds a block_access_list_hash field to the block header, recording every account and storage location accessed during the block's execution along with their post-execution values. With that declaration, clients can read from disk in parallel, verify transactions in parallel, compute the state root in parallel, and even update state without executing at all.
The costs are written into the proposal as well. The access list has to be produced, usually by the block builder after execution, and other nodes must verify its correctness; transaction order, uniqueness, and determinism need to be re-specified, because parallelism presupposes that the locations each transaction depends on have already been declared. The status marked in the EIP-7928 text is peer review, not yet finalized; the Glamsterdam upgrade has brought BAL within the scope of devnet testing, no mainnet activation date is set, progress is tracked on ethereum.org's Glamsterdam roadmap page, and devnet details are maintained by a third-party tracker page. It does not abolish the design origin; it simply adds, for parallel execution, a prior declaration that has to be verified.
Incidentally, the transaction-level access list introduced by EIP-2930 is optional, and the specification does not force a transaction to declare which accounts and slots it will access, so it never became a dependable premise for parallelism. That is why EIP-7928 raises it to the block level and makes the record mandatory.
Counterexamples and boundaries: order exists, but conflicts need not
Explaining the design origin clearly also means stating its boundaries. Order is mandatory, but conflicts are workload-dependent. Transactions such as batch payments, settlements, and airdrop distributions mostly touch only their own recipients' balances, so their read-write sets barely overlap and they could in theory run highly parallel; workloads such as hot AMM pools, liquidations in lending markets, and NFT mints have heavily overlapping read-write sets, and the conflicts themselves eat up the parallel headroom. The same execution model behaves very differently under different workloads, so any discussion of parallel gains has to state the workload profile and the conflict rate.
Another boundary lies in where determinism has to land. Parallel execution does not remove the determinism requirement; it relaxes the object that must be preserved from one actual evaluation order to a class of executions serializable-equivalent to that order. The cost of conflict detection, validation, and selective re-execution lands back on throughput, so parallelism looks more like shrinking the serial segment down to the conflict path than removing serial execution altogether.
One more premise is easy to overlook: however parallel the execution stage is, it must ultimately converge on the same global state tree and the same root. Transactions on the conflict path must be replayed in canonical order, and cross-shard state access requires extra coordination. This is a constraint that later discussions of state sharding cannot get around either.
Division of labor: the origin here, the ceiling and past congestion elsewhere
The only questions to answer here are “why execution must be one transaction at a time, and what it costs.” How high the performance ceiling created by this boundary is, how the congestion episodes since 2017 repeatedly exposed it, and why gas-limit adjustments and EIP-1559 cannot change the execution model belong to the published “Why a Single-Threaded EVM Caps TPS: Congestion History and the Execution Model,” and its figures and reform history are not repeated here. Together the two pieces form a complete chain: first make the origin clear, then measure the ceiling.
Sources
- Ethereum Yellow Paper: the state transition function in Section 2 (equation 1) and the block-level nested definition (equation 2), overall validity in Chapter 4, initial transaction validity in Section 6 (nonce, balance, and EIP-3607), the cumulative gas definition for block receipts (equation 186); the proof space discussion in Appendix D.1.
- EIP-7928: Block-Level Access Lists (Review, created 2025-03-31): the motivation that unknown read-write sets prevent parallelism, the
block_access_list_hashfield, the order and determinism requirements, and the problem that EIP-2930 access lists are not mandatory. - EIP-3607: rejecting transactions from accounts with deployed code.
- ethereum.org: Glamsterdam: the progress statement that BAL has been brought into the scope of Glamsterdam devnet testing, with no mainnet date set.
- reth: stages docs: the division of responsibilities between SenderRecoveryStage and ExecutionStage.
- reth sender_recovery.rs: signature recovery executed in parallel on a rayon thread pool.
- reth_trie_parallel::state_root_task source: the state root task receiving either an execution hook (corresponding to serial execution) or a hashed update stream.
- reth_engine_tree::tree::state_root_strategy: falling back to serial state root computation when a sparse trie task times out or fails.
- reth payload_validator source and payload_processor::prewarm source: sequential EVM execution paired with parallel prewarm pre-execution and a parallel state root task, plus the parallel execution branch when a BAL is present.