The first article in this series described how a large physics model (LPM) gains spatial context by gathering neighboring points and encoding position. The model embeds each sampled surface point, with its position and these geometric features, into a vector, so each point becomes a token. However, this rich set of features is not enough on its own, and that is where attention helps.
Large language models (LLMs) work with tokens in the same way: each word or word fragment of a text is converted into a vector, and attention relates tokens across a whole passage rather than only between adjacent ones (Vaswani et al., 2017). LPMs use attention to capture relationships across a geometry that do not fit inside a small geometric neighborhood. This article covers how point tokens inform one another, and how a model can use information from across a geometry without comparing every point with every other point.
Limitations of proximity-only methods
Many physical fields depend on relationships that reach well beyond a point’s immediate surroundings. In a structure, a load applied at one location is carried along load paths to supports far away. In a conducting part, heat spreads through the material and links regions that are far apart on the surface. In incompressible flow, the pressure at one location depends on conditions throughout the domain. A model that predicts these fields needs a way to relate points that a small geometric neighborhood does not connect.
External aerodynamics provides a clear example. The trailing vortices generated at the wingtips extend well beyond the immediate geometric neighborhood (NASA Glenn). In the wing cases studied in SMART, these vortices are long and coherent, so two separated regions can belong to the same flow structure even when a small-radius neighborhood cannot connect them.
Nearby locations also need not share a state. The upper and lower surfaces of a lifting wing carry different pressure distributions despite sitting only a wing thickness apart, and the two sides of a thin wall in a heat exchanger can sit at very different temperatures. A distance-only rule does not distinguish either pair.
Proximity is useful information but does not decide relevance on its own, so the model has to combine the local detail from Part 1 with information from farther away.
Several methods connect distant regions. Graph networks pass information between neighbors over many steps, multiscale hierarchies work on coarser versions of the geometry, and spectral operators act on the whole domain at once. Attention is another option and is often paired with neighborhood features, as in GeoTransolver, which combines attention with the multi-scale neighbor gathering described in Part 1.
Attention as learned weighting
Attention sets the weight between two tokens from the features each token carries, which include its position and the local geometric context from Part 1. Two points far apart on a geometry can therefore receive a strong weight when their features indicate that they are related, and two adjacent points can receive a weak one when their features differ, as on the upper and lower surfaces of a wing. A distance-only rule applies the same weights regardless of geometry or operating conditions, while attention recomputes them for every input.
The wake is not an input to a geometry-to-field model; the model learns useful representations from geometry, operating conditions, and training examples. Attention exchanges information within that process and does not detect vortices as such or establish physical causality.
Attention can connect spatially separated tokens or concentrate on nearby ones. The widget compares a fixed distance-only rule with an input-dependent, learned weighting.
How attention is computed and what it costs
Attention computes its weights by projecting each token into a query that describes what the token is looking for, a key that describes what it offers for matching, and a value that carries the information passed on. Comparing queries with keys produces the weights that combine the values, and the projections are learned during training.
The direct way to apply attention is self-attention over the whole point cloud, in which the queries, keys, and values all come from the same set of tokens and every point attends to every other point. For N points, that is N × N query–key pairings per attention head, where a head is one of several weight computations an attention layer can run in parallel. The point count grows linearly, but the pairings grow with its square, so doubling the number of points quadruples the work.
At 250,000 points, self-attention evaluates 62.5 billion pairings per head. Stored at 4 bytes per entry, the weights for that single head would occupy 250 GB. Memory-efficient attention kernels avoid storing the full weight matrix, but every pairing still has to be computed, and a model repeats that computation across several heads and layers.
Full self-attention over every point therefore becomes impractical as point clouds grow. To make attention feasible at this scale, the model needs some form of compression: a much smaller set of tokens that summarizes the point cloud, so that attention runs over the small set while the large set stays connected to it. Slicing and cross-attention encoding are two ways to build that compressed representation.
Slicing: group and mix
Slicing, introduced by Transolver, reduces the number of tokens that take part in attention. The model assigns the N point tokens to a much smaller set of M slice tokens, and each slice is a weighted combination of many points. Self-attention then runs over the slices alone, which costs M × M pairings instead of N × N. The updated slices are distributed back to the points with the same assignment weights, so each point receives information gathered from across the geometry.
The assignment is computed from learned features rather than from fixed spatial bins. The illustration below shows the two steps: points contributing to slices, and the slices exchanging information.
Because the assignment is learned, a slice can gather points from beyond a fixed spatial neighborhood and need not correspond to one contiguous patch or to one named vortex.
Cross-attention: encode and refine
In cross-attention, the queries come from one set of tokens and the keys and values from another, so one set reads the other. Cross-attention encoding, as used in Luminary-SMART, builds the compressed representation from M latent tokens, a small set of learned vectors that do not correspond to individual points. The latent tokens act as queries and read the N geometry tokens through cross-attention, which costs N × M pairings instead of N × N; the compression comes from the query set being much smaller than the set it reads. The result is then refined.
Refinement is cross-attention because its queries come from the earlier latent state and its keys and values come from the geometry-enriched state.
What the compact representation must retain
A compact set is cheaper to process but still has to retain the distinctions the prediction depends on, so its size affects accuracy as well as cost.
In SMART, latent capacity is tied to the flow structures and extent of the domain, such as long, coherent wing wakes, rather than to the simulation mesh’s point count. For slices, the aggregation rule matters as well, since combining many points helps only if their distinguishing information is still represented well enough for the task.
The representation, its size, the training data, and the problem together determine what the model can preserve and predict.
From geometric context to information exchange
Part 1 covered how geometry becomes a model input, and this article covered how information is exchanged once those inputs are encoded. Physical fields contain relationships that proximity alone does not describe, attention learns how to combine information across them, and a compact representation keeps that exchange tractable at large point counts.
Self-attention mixes information within a set, and cross-attention lets one set read another. The architecture determines which sets are connected and where the computational cost falls.
To follow the rest of the Demystifying Physics AI series, subscribe to our newsletter.
- Points
- Pairings per head
- Weights if stored (FP32)