What Is PTX and How Does It Relate to SASS in GPU Execution?

PTX (Parallel Thread Execution) is the intermediate representation a GPU compiler emits for NVIDIA hardware: it is a virtual instruction set, not the machine code that actually runs. SASS is the architecture-specific machine code the driver's assembler produces from PTX for a particular GPU generation. If you want to reason about compiled GPU behavior, the useful mental model is a two-stage lowering path — high-level code to PTX, then PTX to SASS — where PTX stays portable and SASS does not.

The lowering path, stage by stage

  1. High-level GPU code (CUDA C++, or another language with an NVIDIA backend) is compiled by a front-end compiler such as NVCC.
  2. PTX is emitted as the compiler's target-independent output. It is a defined virtual ISA with its own types, registers, and instruction set.
  3. The driver's assembler (ptxas) translates PTX into SASS for the specific GPU architecture in use.
  4. SASS is what the hardware executes. It encodes real registers, real instruction encodings, and scheduling decisions tied to that architecture.

The key consequence: PTX is a contract, not a final artifact. The same PTX can be assembled into different SASS for different GPU generations.

Why PTX is portable and SASS is not

Property PTX SASS
Level Virtual ISA / intermediate representation Native machine code
Portability Portable across GPU generations Specific to one architecture family
Produced by Front-end compiler (e.g. NVCC) Driver assembler (ptxas)
Stability Documented, relatively stable Undocumented, varies by architecture
Role Forward-compatible distribution format What the hardware actually runs

Because PTX is portable, shipping PTX lets a future driver re-assemble it for newer hardware. Because SASS is architecture-specific, a binary compiled for one generation is not guaranteed to run on another. That asymmetry is the whole reason the two-stage design exists.

How the distinction affects performance analysis and debugging

The two levels answer different questions:

  • PTX tells you what the compiler intended: which operations were generated, how memory was addressed, whether certain optimizations were applied at the IR level.
  • SASS tells you what the hardware will actually do: the real instruction mix, register allocation, and scheduling. Performance characteristics such as instruction counts and stall behavior live here.

Practical implications:

  • If you inspect only PTX, you can miss decisions made during assembly — register pressure, instruction selection, and scheduling that change runtime behavior.
  • If you inspect only SASS, you lose the higher-level intent and the mapping back to source constructs.
  • For debugging, PTX is the more stable reference; for performance, SASS is closer to ground truth.

A concrete example: a loop that looks efficient in PTX may assemble into a SASS sequence with extra instructions or different register usage on one architecture versus another. The PTX is identical; the executed code is not.

Where tools like Gestell fit

Gestell is focused on compiled GPU execution analysis — building tools to understand and advance the frontier of GPU performance, and studying PTX, SASS, compiler lowering, and GPU execution behavior. That places it at exactly the boundary this article describes: the point where PTX is lowered into SASS and where execution behavior is determined. If your goal is to reason about why compiled GPU code behaves as it does, the PTX-to-SASS transition is the layer to examine, and tooling aimed at that layer is where the analysis happens.

What to take away

  • PTX is an intermediate representation, not the final machine code.
  • SASS is the architecture-specific machine code that actually executes.
  • The path is: high-level code → PTX → SASS, with the driver's assembler performing the second step.
  • PTX is portable across GPU generations; SASS is not.
  • For performance and debugging, use PTX to understand intent and SASS to understand execution.
gestell.ai
Gestell studies PTX, SASS, compiler lowering, and GPU execution behavior.