a graphics tutorial · from your very first pixel

zimr, from scratch

You have never written a shader. You are not sure what a GPU does, exactly. This page starts there and ends with you reading — and understanding — a fragment shader that sphere-traces glowing metaballs and makes them collide with ordinary cubes.

We learn through zimr: a small engine written entirely in one language, Zig, where shaders are just functions, the GPU is set up for you, and — the idea everything else hangs on — the same function you write for the GPU also runs on the CPU. That turns out to be the best teaching tool in graphics, because you can read a shader like normal code before you ever trust it to a chip you can't step through.

Lesson 01

What a GPU is really doing

Goal: replace the word "GPU" with a mental picture you can actually reason about.

Here is the whole idea. Your screen is a grid of pixels — say two million of them. To draw a frame, some code has to decide on a color for every one of those two million pixels, sixty times a second. That is a hundred and twenty million color decisions per second, and that is a slow frame.

A CPU is a few very clever workers who do tasks one after another, extremely fast. A GPU is the opposite bargain: thousands of simpler workers who all do the same task at the same time, each on a different piece of data. Two million pixels is exactly that shape of problem — the same "what color are you?" question, asked two million times, with no pixel needing to know the answer for any other. So we hand it to the crowd.

The small program that the crowd runs — the recipe every worker follows to answer "what color are you?" — is called a shader. That is the entire mystery. A shader is a function. It runs once per pixel (or once per corner of a triangle), millions of copies at once, and its job is to return a color or a position. When people say a shader "runs on the GPU," they mean: the crowd all runs this one function, each on their own pixel.

Two shaders, two jobs Almost everything on screen is triangles. A vertex shader runs once per triangle corner and decides where that corner lands on screen. Then the hardware fills in the triangle and runs a fragment shader once per pixel inside it, to decide what color that pixel is. Position, then color. Hold onto that pair — the rest of this page is just those two functions getting more interesting.

The catch, historically, is that these functions had to be written in a separate little language (GLSL, or on the web, WGSL), shipped as text, compiled by the driver at runtime, and debugged by staring at a black screen because you can't put a breakpoint inside a crowd of two million. zimr's whole bet is to make that catch go away.

Lesson 02

The smallest zimr program

Goal: see that an app is a struct of state plus three functions — and there is no main loop to write.

Before any graphics, meet the shape of a program. In zimr an app is a piece of state (a plain struct — your variables) plus a lifecycle of three functions. The host owns the browser's frame loop and calls your update sixty times a second. You never write while (true).

a complete app — bounce a circleconst z = @import("zimr");  // the engine: draw*, input, textures…
const zm = @import("zm");   // the math: Vec, Mat, dot, normalize…

// Your state. An ordinary struct. No globals, no singletons.
const State = struct {
    pos: zm.Vec2 = .{ 200, 150 },
    vel: zm.Vec2 = .{ 3, 2 },
};

pub const app: z.AppSpec(State) = .{
    .config = .{ .window = .{ .title = "bounce", .width = 800, .height = 450 } },
    .init = initState,   // once, at startup → returns State
    .update = update,    // every frame → input + draw
};

fn initState(gpa: std.mem.Allocator, f: *z.Frame) !State {
    return .{}; // defaults are fine here
}

fn update(f: *z.Frame, s: *State) void {
    s.pos += s.vel; // Vec2 arithmetic is just +  (see Lesson 06)
    if (s.pos[0] < 20 or s.pos[0] > 780) s.vel[0] = -s.vel[0];
    if (s.pos[1] < 20 or s.pos[1] > 430) s.vel[1] = -s.vel[1];

    z.beginDrawing(f.gl);
    z.clearBackground(f.gl, z.Color.ray_white);
    z.drawCircleV(f.gl, s.pos, 20, z.Color.maroon);
    z.endDrawing(f.gl);
}

That is a real, running program. Three things to notice, because they are true of every zimr app:

Lesson 03

The frame, and drawing in 2D

Goal: understand "immediate mode" and why adding draw calls is cheap.

When you call drawCircleV, nothing is sent to the GPU yet. zimr is immediate mode: each draw call quietly appends triangles to a growing list (a circle is a fan of thin triangles; a line is two). Only at endDrawing does that whole list get handed to the GPU in a single go — one render pass.

That's why a scene with two hundred shapes is not two hundred trips to the GPU; it is one trip carrying more triangles. The cost of a draw call is vertices, not conversations with the hardware. Reading input works the same relaxed way — the frame holds a snapshot and you ask it questions:

input is polled, not pushed at youfn update(f: *z.Frame, s: *State) void {
    if (z.isKeyPressed(f.input, .space)) s.paused = !s.paused; // once, on the down-edge
    if (z.isKeyDown(f.input, .right))  s.pos[0] += 4;            // every frame it's held

    const touches: i32 = z.getTouchPointCount(f.input);        // phones report multitouch
    const dt: f32 = @floatCast(f.time.delta_time);            // seconds since last frame
    ...
}
Why every zimr demo works on a phone Because input is just a snapshot you interrogate, the engine can offer a gesture layer on top — taps, holds, swipes, pinches — and every 3D demo in the tree shares one idiom: one finger drags to orbit the camera, two fingers pinch to zoom. The drag is gated on whether the UI wants the touch first, so grabbing a slider never also spins the world.
Lesson 04

Your first shader is a function

Goal: write a fragment shader and realize it is just Zig.

So far the engine has been coloring pixels for us. Now we take the wheel. Remember the fragment shader: it runs once per pixel and returns a color. In zimr you write it as an ordinary Zig function named shaderMain. Here is one that paints a smooth vertical gradient — no textures, no tricks, just "turn my vertical position into a color":

runs on GPU runs on CPU comptime
a fragment shader that is a gradientconst zm = @import("zm");
const Vec2 = zm.Vec2;

pub fn shaderMain(io_in: Io) Out {
    var out: Out = undefined;

    // frag_uv is (0,0) at one corner, (1,1) at the opposite corner.
    // It was handed to us per-pixel, smoothly interpolated. (Lesson 06.)
    const uv: Vec2 = io_in.frag_uv;

    // Bottom is warm, top is cool. That's it — a color per pixel.
    out.final_color = .{ uv[1], 0.3, 1.0 - uv[1], 1.0 };
    return out;
}

There is no new language here. Vec2 is a Zig vector; uv[1] is its second component; the return value is four floats — red, green, blue, alpha. If you can read Zig, you can read a shader, because the shader is Zig. The build system, not you, turns this into the WGSL that the browser's GPU actually consumes. You never see that WGSL, the same way you never see the assembly your C compiler emits.

your shaderMain (Zig) │ zig build-obj -target spirv32 ▼ SPIR-V (a portable GPU bytecode) │ spv2wgsl (a translator, itself written in Zig) ▼ WGSL → the browser → the GPU

That middle tool, spv2wgsl, is doing real work — GPU bytecode and WGSL disagree about how loops and branches are even shaped, so it rebuilds the control flow — but from where you sit, it's invisible. You wrote a function; a color appeared on two million pixels.

Lesson 05

One function, three homes

Goal: the central idea of the whole engine — and why it makes you a better graphics programmer.

Look again at that gradient shader. Nothing in it is GPU-specific. It takes some numbers in, does arithmetic, returns numbers out. So here is zimr's core move: that exact same function is compiled to run in three different places.

▦
the GPU
→ SPIR-V → WGSL. Two million copies, in parallel, at full speed. The real thing.
▤
the CPU
zimr has a software rasterizer — it calls the same shaderMain per pixel, in plain Zig you can step through.
◈
comptime
some shaders are even run by the Zig compiler at build time, baking their output straight into the binary.

Why does this matter to you, on your first day? Because it dissolves the worst thing about learning graphics. Normally a shader is a black box: it runs somewhere you can't reach, and when it's wrong you get a black screen and no error. With zimr, before a GPU ever sees your shader, the engine can run the identical function on the CPU and show you the picture — and there you can add a print statement, set a breakpoint, and watch a single pixel's math step by step.

The aha A shader stops being a mysterious thing that runs "on the graphics card" and becomes what it always was underneath: a function from numbers to numbers. The GPU is just the fastest place to run it. Every stylized demo further down this page is checked on the CPU first — pixel-for-pixel — and only then trusted to the GPU. You get to learn the same way.

One source of truth, three execution sites, decided by the build and not by you. Keep this picture; the "runs on…" chips above each code block from here on tell you which homes that particular snippet lives in.

Lesson 06

Vertices, varyings, and the schema

Goal: understand where frag_uv came from, and meet the typed schema that ties the two shaders together.

In Lesson 04 the fragment shader received io_in.frag_uv. Where did that come from? It was produced by the vertex shader and smoothly blended across the triangle on the way to each pixel. A value that the vertex stage computes at the corners and the hardware interpolates for every pixel in between is called a varying. Position in, varying out; that interpolation is free hardware magic and it is where a surprising amount of graphics beauty comes from.

For the two shaders to agree on what's being passed between them, zimr asks you to declare it — as plain Zig structs, in a sibling file. A shader is two small files: the body (foo_fs.zig) and its schema (foo_fs_io.zig). Here is a real one, the vertex stage that every lit 3D material in the engine shares:

src/shaders/gbuffer_vs.zig — a real vertex shaderpub fn shaderMain(io_in: Io) Out {
    var out: Out = undefined;

    // WHERE this corner lands on screen: multiply the vertex by the
    // model-view-projection matrix. This one line is 3D graphics.
    out.position = mulMatPoint(io_in.u.mvp, io_in.vertex_position);

    // Two facts the fragment stage will want, computed here and passed
    // along as varyings: the world-space position and the world normal.
    const world: Vec = mulMatPoint(io_in.u.model, io_in.vertex_position);
    out.frag_world_pos = Vec3{ world[0], world[1], world[2] };

    const n: Vec3 = io_in.vertex_normal;
    const wn: Vec = mulMatVec(io_in.u.normal_matrix, vec4(n[0], n[1], n[2], 0.0));
    out.frag_world_normal = Vec3{ wn[0], wn[1], wn[2] };

    return out;
}

That single line — out.position = mulMatPoint(mvp, vertex_position) — is the beating heart of all 3D. A matrix (zm.Mat, a 4×4 grid of numbers) encodes "move the camera here, point it there, apply perspective." Multiplying a point by it projects that point onto the screen. You will build these matrices with helpers like lookAtRh (aim a camera) and perspectiveFovRh (add perspective) and mostly not think about their insides — but now you know what the vertex shader is for.

Why the math reads like math zm.Vec is literally @Vector(4, f32) — a native Zig vector. So a + b, a * scalar, and uv[1] just work, with no library ceremony. The engine leans on this hard: everything from a physics step to a fragment shader speaks the same zm vocabulary, which is exactly what lets one function run in all three homes.
Lesson 07

The uniform block, from both sides

Goal: get data from your CPU program into a running shader — without the two sides ever disagreeing about the layout.

Your shader needs values that are the same for every pixel this frame: the light's direction, the camera position, a color, the current time. That bundle is called a uniform block (uniform = same for all pixels). You declare it once, as a Zig struct, in the schema file:

src/shaders/fog_fs_io.zig — the schema (condensed)pub const Ubo = extern struct {
    base_color: Vec,
    view_pos:   Vec,   // where the camera is, in world space
    light_dir:  Vec,   // which way the light points
    fog_color:  Vec,   // what the distance dissolves into
    params:     Vec,   // { fog_density, 0, 0, 0 }
};

Now the payoff of declaring it in Zig. On the host — your normal CPU-side program — you ask for that exact same type back, fill it in, and upload it:

host side the same struct, in the shader
on the host — fill the block and send it// Take the shader's uniform type straight from the shader's Io.
const FsUbo = @FieldType(fog_fs.Io, "u");

var ubo: FsUbo = .{
    .base_color = .{ 0.8, 0.5, 0.3, 1 },
    .light_dir  = .{ 0.4, 0.8, 0.4, 0 },
    .fog_color  = .{ 0.1, 0.1, 0.12, 1 },
    .params     = .{ s.density, 0, 0, 0 },
    ...
};
z.wgpu.queueWriteBuffer(f.gpu.queue, s.fs_ubo, 0, std.mem.asBytes(&ubo));
The aha The host and the shader are reading the same struct definition. There is no hand-written list of "byte 0 is the color, byte 16 is the light direction" that the two sides have to keep in sync by hand — the classic way a shader silently renders garbage. One Zig type is the single source of truth for the memory layout, used by both sides. If you add a field, both sides see it, or neither compiles.

Textures work the same tidy way: a sampler is a field (texture0: shader.Sampler2D(...)) and reading it is a function call in the shader (io_in.texture0(uv)). Declare it, then call it.

Lesson 08

Light: the whole trick

Goal: understand the one equation that makes 3D look 3D.

Here is the single most useful fact in shading. A surface looks bright when it faces the light and dark when it faces away. "How much a surface faces the light" has an exact measure: the dot product of the surface's normal (the little arrow pointing straight out of it) and the direction to the light. Facing dead-on gives 1; edge-on gives 0; facing away goes negative, which we clamp to 0. That's it. That's Lambert lighting, and it is most of what your eye reads as shape.

the four lines under almost every lit surfaceconst normal: Vec3 = normalize(io_in.frag_world_normal);      // the surface's arrow
const light_dir: Vec3 = normalize(/* toward the light */);
const n_dot_l: f32 = @max(dot(normal, light_dir), 0.0);   // how much it faces the light
const lit: Vec3 = base_color * @as(Vec3, @splat(n_dot_l)); // dim it by that amount

Notice dot and normalize are just zm functions — the same ones a CPU simulation would call. A pixel that faces the sun keeps its full color; one that faces away goes dark. Wrap a sphere in this and it suddenly reads as round. Everything fancier — specular highlights, fog, shadows — is a refinement bolted onto this skeleton. The distance-fog material in the tree, for instance, does exactly this and then mixes the result toward a fog color by distance, using raylib's exponential curve.

Lesson 09

Cel shading — one line changes everything

Goal: feel how a tiny change to the lighting math becomes a whole art style. This is the fun part.

Take that smooth Lambert term — n_dot_l, a number sliding continuously from 0 to 1 across a curved surface — and do one violent thing to it: snap it to a few discrete steps. Smooth becomes stair-stepped, and a rounded object suddenly looks hand-inked, like a comic book. This is cel shading, and it is genuinely almost one line. Here is the real shader from the tree:

GPU CPU (verified here first) comptime
src/shaders/cel_fs.zig — toon shading, the real thingpub fn shaderMain(io_in: Io) Out {
    var out: Out = undefined;

    const normal: Vec3 = normalize(io_in.frag_world_normal);
    const light_dir: Vec3 = normalize(Vec3{
        io_in.u.light_dir[0], io_in.u.light_dir[1], io_in.u.light_dir[2],
    });
    const n_dot_l: f32 = @max(dot(normal, light_dir), 0.0);

    // THE PUNCHLINE: snap the smooth 0..1 falloff into `bands` flat
    // plateaus. floor() is the whole trick — it turns a gradient into steps.
    const bands: f32 = @max(io_in.u.params[0], 2.0);
    const quantized: f32 =
        @min(@floor(n_dot_l * bands), bands - 1.0) / (bands - 1.0);

    const lighting: f32 = @min(ambient_floor + quantized, 1.0);
    out.final_color = .{
        io_in.u.base_color[0] * lighting,
        io_in.u.base_color[1] * lighting,
        io_in.u.base_color[2] * lighting,
        1.0,
    };
    return out;
}

Read that quantized line slowly, because it's a perfect example of what shader programming actually is. n_dot_l is between 0 and 1. Multiply by bands (say 4) to get 0–4, floor it to get a whole step (0, 1, 2, or 3), divide back down. A continuous ramp becomes four flat regions with hard borders between them. One @floor is the entire art style.

A scar in the comments The real file divides by bands - 1, not bands, so the brightest plateau reaches a full 1.0. The comment next to it says: "divide by bands and nothing ever hits white — the whole render reads underexposed; ask us how we know." Shaders are full of these off-by-a-hair traps, which is exactly why being able to run the thing on the CPU and inspect one pixel's arithmetic is worth so much.
Lesson 10

Sharing a vertex stage

Goal: see the design idea the whole shader library is organized around — and why adding a new look is one file.

You may have noticed the cel shader and the fog shader both read frag_world_normal and frag_world_pos. That is not a coincidence — they share a vertex shader, the gbuffer_vs from Lesson 06. Because varyings are a named Zig struct, a fragment shader that wants those inputs just aliases them:

a fragment schema borrowing a shared vertex stage// src/shaders/fog_fs_io.zig
const common = @import("gbuffer_common_io.zig");
pub const Inputs = common.Interp; // ← "I take exactly what gbuffer_vs emits"

That one line means fog, cel shading, terrain banding, the maze walls in the voxel demos — every world-lit material — all pair with the same vertex shader. Stage-to-stage compatibility is guaranteed by the type system, not by matching up numbered slots by hand (the traditional, error-prone way). So the engine has grown two universal vertex stages that most materials just plug into:

shared vertex stage what it emits who plugs in
gbuffer_vs clip position + world position + world normal every lit surface — fog, cel, terrain, maze walls. A new material is one fragment file.
deferred_shading_vs a full-screen quad and its UV every screen-space effect — wave distortion, bloom, the raymarcher below. A new effect is one fragment file.
The aha "Add a new visual style" collapses to "write one small fragment function." That is why the effects gallery in the engine is dozens of looks deep without dozens of pipelines — they nearly all borrow one of these two vertex stages and differ only in the per-pixel math you now know how to read.
Lesson 11

How WebGPU actually gets set up

Goal: peek under beginDrawing and see the real WebGPU objects — so the word stops being scary.

For 2D and the immediate-mode 3D batch, zimr does all of this for you. But when a demo wants its own shader (the heightmap, the metaballs), it drops one level down to f.gpu and talks to WebGPU directly. WebGPU is the modern browser graphics API; here are its four nouns, in the order you meet them.

nounwhat it is
Buffera block of GPU memory — your vertices live here, your uniform block lives here.
Shader moduleyour compiled shaderMain, loaded onto the device.
Pipelinethe recipe tying it together: this vertex layout, this vertex + fragment shader, draw triangles, test depth this way.
Bind groupthe wiring that says "uniform buffer #2 goes to the shader's u." Built from the layout the schema generated.

Setup happens once, in init. This is condensed from the real heightmap demo — a generated terrain mesh drawn with a custom material — and it is representative of every custom-pipeline demo in the tree:

host — init, runs once
creating a pipeline (condensed from examples/heightmap)// 1. GPU memory for the mesh vertices and the uniform blocks.
const vbo = z.wgpu.createBuffer(device, .{ .size = cap, .usage = .{ .vertex = true, .copy_dst = true } });
const vs_ubo = z.wgpu.createBuffer(device, .{ .size = @sizeOf(VsUbo), .usage = .{ .uniform = true, .copy_dst = true } });

// 2. Load the two shaders (compiled from Zig → WGSL by the build).
const vs_mod = z.wgpu.createShaderModuleWgsl(device, gbuffer_vs_wgsl, "gbuffer_vs");
const fs_mod = z.wgpu.createShaderModuleWgsl(device, terrain_fs_wgsl, "terrain_fs");

// 3. The pipeline: vertex layout + shaders + how to draw.
const combo = z.gpu.StateCombo.fromParts(.triangle_list, .none, .less, .back, .rgba8_unorm, .depth24_plus, 1);
const pipeline = z.wgpu.createRenderPipeline(device, layout, vs_mod, fs_mod, pipe_blob, "hm_pipe");

Then each frame you fill the uniform buffer (Lesson 07), open a pass, set the pipeline and bind groups, and draw. It is more ceremony than drawCircleV, but every piece is a noun from the table above, and — crucially — the scary layout-matching part is gone, because the vertex layout and the bind-group layout were both generated from your Zig schema.

A real bug this design catches WebGPU runs every buffer upload before the frame draws. So writing the same uniform buffer twice in one frame means only the last write survives — a classic way to get a corrupted or "duplicated" look. zimr ships a headless smoke test that scans each frame's uploads and fails the build if it catches the same buffer written twice. It recently caught exactly this in the shared 3D batch: a leftover startup write to the camera buffer, harmless but flagged, and removed. The tooling watches the seams so you don't have to.
Lesson 12

The showpiece: raymarched metaballs

Goal: read a genuinely advanced shader and find that you understand every idea in it.

Everything so far has drawn triangles. This one doesn't. A raymarching shader runs on a plain full-screen quad, and for each pixel it shoots a ray out into an imaginary world and marches along it in steps until it bumps into a surface — a surface defined not by triangles but by a math formula (a "signed distance function"). It is a whole 3D scene conjured purely inside the fragment shader.

The one in the tree, src/shaders/hybrid_raymarch_fs.zig, marches three metaballs that melt into each other (a smooth-minimum union) floating over a checkerboard floor. Its header comment describes the scene and then the clever bit:

src/shaders/hybrid_raymarch_fs.zig — the header//! The scene: three metaballs orbiting each other (polynomial
//! smooth-min union) over an analytically-intersected checkerboard
//! floor.  Shading is a tetrahedron-gradient normal, Lambert from a
//! fixed sun, a rim term on the blobs, and distance fade on the floor.
//!
//! The hybrid trick: on a hit, the world-space point is projected
//! through the SAME view-projection the raster pass uses and
//! clip.z / clip.w goes out through the FragDepth builtin — so
//! rasterized cubes and these marched blobs depth-test against each
//! other with zero coordination.

Look at how much of that you can now decode. "Lambert from a fixed sun" — that's Lesson 08, the dot product. "A normal" — the surface arrow the lighting needs; here it's recovered from the distance formula instead of read from a mesh. "Projected through the same view-projection the raster pass uses" — that's the vertex-shader matrix from Lesson 06, reused so that these formula-defined blobs and ordinary triangle cubes correctly hide each other. Different technique, same handful of ideas.

And it was debugged on the CPU A raymarcher is the last thing you'd want to debug blind on a GPU — it's a loop with branches, per pixel. Because this shaderMain is plain Zig over zm, it ran first in the software rasterizer, where a wrong ray is an ordinary Zig bug you can print and step. Lesson 05's promise, cashed in on the hardest shader in the tree.
Lesson 13

The GPU as a calculator (compute)

Goal: see that the GPU's crowd-of-workers is useful even when you're not drawing.

A shader doesn't have to make pixels. If you have a big array and the same simple operation to do to every element, the GPU's thousands of workers will chew through it far faster than the CPU. That's a compute shader, and zimr writes them in the same one-source style. Here is the smallest possible one — "double every number" — and notice the comment:

GPU (thousands at once) CPU (identical loop)
examples/compute_smoke/double_it.zig/// The kernel. Identical on CPU and GPU; buffers reached via `b_data`.
pub fn double(c: k.Ctx(@This())) void {
    if (c.id >= c.params.count) return; // which element am I? (my worker id)
    b_data[c.id] = b_data[c.id] * 2.0;
}

Each worker gets an id — "you handle element 7" — checks it's in range, and does its one small job. Launch a thousand workers and a thousand elements double at once. The same kernel runs as a plain CPU loop when you want to verify it. zimr uses this for particle systems, a fluid simulation, and a smoke sim — all the same pattern: a small function, run across a huge array, checkable on the CPU first.

Lesson 14

Where to go next

Goal: turn understanding into a first change of your own.

You started not knowing what a GPU does. You can now read a lit material, a toon shader, and the header of a raymarcher, and see the same small set of ideas — position by a matrix, color by a dot product, data shared through a typed struct — dressed up in different clothes. That is most of graphics.

The shortest path from here to a change you made yourself:

Every shader you read here lives in src/shaders/; every demo lives in examples/. They are all just Zig — which means the compiler, the debugger, and the print statement you already know work on the whole stack, GPU and CPU alike. Go change a number and watch the pixels move.