Normally a frame draws straight to the screen — the swap-chain backbuffer. Render-to-texture (RTT, also "render textures" or "framebuffer objects") inserts a detour: you draw into an offscreen image instead, and then that image becomes an ordinary texture you can sample, scale, tint, feed to a shader, or read back to the CPU.
It's a two-phase move. Phase one: point your draw calls at the offscreen target and render a scene into it. Phase two: go back to drawing on the screen, and use the thing you just rendered as input. The six-panel demo this tutorial accompanies is the clearest possible illustration — one animated scene is rendered once into a 256×256 texture, then that single texture is stamped onto the screen six times at different tints. Render once, reuse many.
RTT is the substrate under a surprising amount of graphics work:
copy_src, so its pixels can be pulled back to the CPU for screenshots, hit-testing, or image export.Five calls, shaped to mirror raylib:
z.loadRenderTexture(gl, w, h) → a RenderTexture (color texture + view + sampler, optional depth).z.beginTextureMode(gl, rt) — redirect drawing into rt.z.endTextureMode(gl) — stop, return to the screen.rt.asTexture() — view the result as a WgpuTexture for drawTextureRec.z.unloadRenderTexture(gl, &rt) — free the GPU resources.The usage rhythm — note that the whole thing happens inside a normal beginDrawing/endDrawing frame:
z.beginDrawing(f.gl);
z.clearBackground(f.gl, dark);
z.beginTextureMode(f.gl, rt, dark); // phase one: draw into the texture (dark clears; null accumulates)
z.drawCircleV(f.gl, center, r, white);
z.endTextureMode(f.gl);
z.drawTextureRec( // phase two: use it on the screen
f.gl, rt.asTexture(),
0, 0, 1, 1, // source UVs (normalized, whole texture)
dx, dy, dw, dh, // destination rect (screen px)
tint,
);
z.endDrawing(f.gl);
Inside one frame there is a single command encoder and, normally, a single render pass writing to the backbuffer. beginTextureMode performs a mid-frame pass switch: it flushes the pending 2D batch, ends the current pass, and opens a new pass on the same encoder aimed at the render texture's view — with the 2D ortho reprogrammed to the texture's pixel size, so your coordinates map to the texture rather than the window.
endTextureMode reverses it: flush, end the texture pass, reopen the backbuffer pass with a load (not a clear — so whatever you drew to the screen before the bracket survives), and restore the viewport ortho. Because both passes share the frame's one encoder, the GPU executes them in recorded order at submit time: the texture pass first, the screen pass second. That ordering is exactly the dependency we need — the texture is finished before the screen samples it.
pub fn loadRenderTexture(
gl: *WgpuGl,
width: i32,
height: i32,
) WgpuRenderTexture {
const app: *App = appOf(gl);
return WgpuRenderTexture.create(app.gpu_frame.device, .{
.width = @intCast(width),
.height = @intCast(height),
.label = "render_texture",
});
}
/// Free a render texture's GPU resources (raylib `unloadRenderTexture`).
pub fn unloadRenderTexture(gl: *WgpuGl, rt: *WgpuRenderTexture) void {
_ = gl;
rt.deinit();
}
/// Redirect 2D drawing to the offscreen `rt`. Flushes the current pass, opens a fresh
/// pass targeting the render texture, and sets the 2D ortho to the render texture's size.
/// `clear` is explicit: a color clears the texture on entry; `null` PRESERVES last frame's
/// contents (load) so you can accumulate — fade + draw for trails. Pair with
/// `endTextureMode`. Call between begin/endDrawing.
pub fn beginTextureMode(
gl: *WgpuGl,
rt: WgpuRenderTexture,
clear: ?types.Color,
) void {
const app: *App = appOf(gl);
Backend.flushBatch(&app.pass);
Backend.endRenderPass(&app.pass);
app.pass = Backend.beginRenderPass(app.gpu_frame.encoder, .{
.color_view = rt.color_view,
.clear = if (clear) |col| .{ .r = col.r, .g = col.g, .b = col.b, .a = col.a } else null,
.depth_view = rt.depth_view,
});
const vp: [16]f32 = renderer_2d.orthoTopLeft(
@floatFromInt(@max(rt.width, 1)),
@floatFromInt(@max(rt.height, 1)),
);
app.renderer_2d.?.updatePerFrame(&app.gpu_frame, .{ .view_projection = vp });
app.renderer_2d.?.bindForPass(&app.pass);
app.pass.batch = &app.renderer_2d.?.shapes_batch;
}
/// End offscreen rendering: flush, close the render-texture pass, reopen the backbuffer
/// pass PRESERVING what was drawn before `beginTextureMode` (load, not clear), and
/// restore the 2D ortho to the live viewport.
pub fn endTextureMode(gl: *WgpuGl) void {
const app: *App = appOf(gl);
Backend.flushBatch(&app.pass);
Backend.endRenderPass(&app.pass);
app.pass = Backend.beginRenderPass(app.gpu_frame.encoder, .{
.color_view = app.gpu_frame.surface_view,
.clear = null,
.depth_view = app.gpu_frame.depth_view,
});
const size: wgpu.SurfaceSize = wgpu.getSurfaceCssSize(app.gpu_frame.surface);
const css_w: f32 = @floatFromInt(@max(size.width, 1));
const css_h: f32 = @floatFromInt(@max(size.height, 1));
const vp: [16]f32 = switch (app.config.window.scale_mode) {
.responsive => renderer_2d.orthoTopLeft(css_w, css_h),
.fit => fitOrtho(
@floatFromInt(app.config.window.width),
@floatFromInt(app.config.window.height),
css_w,
css_h,
The display side needs the result as something drawTextureRec can bind, so the render texture carries its own sampler and views itself as a WgpuTexture:
pub fn asTexture(self: WgpuRenderTexture) WgpuTexture {
return .{
.handle = self.color,
.view = self.color_view,
.sampler = self.sampler,
.width = self.width,
.height = self.height,
.format = self.format,
};
}
The API we shipped is one point on a spectrum running from implicit and stateful to explicit and threaded. Here is the whole spectrum, roughly most-magical to least, with what each buys and what it costs.
1. Implicit mode — what we chose. A hidden "current target" lives in the app; bracketing with begin/end redirects every subsequent draw…(gl, …) call to the texture, then back. This is raylib's BeginTextureMode/EndTextureMode, verbatim.
z.beginTextureMode(f.gl, rt, clear);
z.drawCircle(f.gl, ...); // goes to the texture
z.endTextureMode(f.gl);
z.drawTextureRec(f.gl, rt.asTexture(), ...); // goes to the screen
Buys: raylib examples port one-to-one — the call shape is identical, so a port is mechanical. And the "magic" is not something we invented: raylib hides exactly this, binding an OpenGL framebuffer object inside BeginTextureMode. We are faithful to the very thing we're porting. Costs: it is stateful — an unbalanced begin/end corrupts the frame — and the pass switch is invisible at the call site.
2. Target-in-handle. Instead of a global, beginRenderTexture returns a second *WgpuGl bound to the texture; you draw through that handle.
const rt_gl = z.beginRenderTexture(f.gl, rt);
z.drawCircle(rt_gl, ...); // the target lives in the handle, not a global
z.endRenderTexture(rt_gl);
Buys: the target is visible in the code — which handle you pass — and there is no hidden global. Still immediate-mode. Costs: two handles to keep straight, ending order still matters, and it no longer matches raylib, so every port needs a small rewrite.
3. Scoped callback. You hand the engine a draw function; it brackets begin/end around it for you.
z.withRenderTexture(f.gl, rt, &ctx, drawSceneFn);
Buys: you cannot forget to end — the scope is enforced. Costs: Zig has no closures, so you thread a ctx struct plus a function pointer by hand, which is clunky for the common case, and it diverges from raylib's call shape.
4. Explicit pass objects — the WebGPU shape. A render pass is a first-class value; draws are methods on it.
var p = z.beginRenderPass(.{ .target = rt, .clear = black });
p.drawCircle(...);
p.end();
Buys: no hidden state at all; multiple passes compose naturally; this is closest to how wgpu actually works underneath. Costs: every draw becomes p.draw…() rather than draw…(gl), which breaks not just raylib parity but the immediate-mode feel of the entire rest of zimr; it is more verbose; and batching has to be tracked per-pass.
5. Per-call target — no modes at all. Each draw names its destination; there is nothing to enter or leave.
z.drawCircleTo(rt, ...); // to the texture
z.drawCircle(...); // to the screen
Buys: maximally explicit and stateless — there is no "current target" to get wrong. Costs: it doubles the draw surface (a …To variant per primitive), batching must key on the target, and it is the furthest thing from raylib.
6. Render graph. Record draw commands into lists, declare which textures each pass reads and writes, and let a scheduler order the passes and insert the dependencies.
Buys: this is how large engines manage many interdependent passes — shadow maps feeding lighting feeding post-processing. Costs: a completely different programming model, overkill for an immediate-mode 2D library, and a poor fit for "port the raylib examples."
The mission decides it. zimr exists to run raylib's example corpus on WebGPU, and raylib's render-texture API is the implicit mode. Matching it means every example that uses a render texture ports with zero adaptation; choosing any of #2–#5 would tax each of those ports with a rewrite, in exchange for a safety-and-clarity win that raylib itself declined to take. So #1 is the right call for a port — even though #4 would likely be the better choice for an engine designed from a blank page.
The part that makes this comfortable to commit: we did not give up the explicit option, we just declined to put it on top. beginTextureMode is a thin wrapper over primitives that already exist and are already explicit — WgpuBackend.beginRenderPass(encoder, .{ .color_view, .clear, .depth_view }) hands back a real pass object, and render_pass.setPipeline / draw / end drive it directly. The implicit, raylib-facing surface sits on top of the explicit pass layer. The day we want #4 — multi-target output, custom load/store ops, a pass that feeds another pass — the machinery is already there to expose. Shipping #1 does not lock it out; it just keeps it off the path the examples walk.
beginTextureMode leaves the frame pointed at the wrong target and corrupts output. Treat the bracket as inseparable.beginDrawing and endDrawing — it manipulates the live pass and encoder, which only exist during a frame.drawTextureRec samples v=0 at the top — so the result is upright, with no negative-height flip of the kind raylib needs for its bottom-up FBOs. The six-panel demo confirms it: the scene is right-side-up.beginTextureMode takes a clear: ?Color — pass a color to clear the texture on entry (the "redraw every frame" pattern), or pass null to LOAD, preserving last frame's pixels so you can fade and accumulate. A motion trail is then just: enter with null, draw a faint translucent-black rectangle over the whole texture to decay the old image, draw the new dots, exit. (See the companion trails demo.)The complete source behind the six tinted panels — render the scene into the texture once per frame, then composite it across a viewport-relative grid:
//! wgpu_render_texture — render-to-texture demo for the new offscreen API. An animated
//! scene (spinning rectangles + a pulsing circle) is drawn INTO an offscreen render
//! texture via `beginTextureMode`/`endTextureMode`, then that ONE texture is composited
//! to the screen as a tinted grid with `drawTextureRec` — "render once, reuse many."
//! Exercises loadRenderTexture / texture-mode / asTexture. Viewport-relative.
const std = @import("std");
const z = @import("zimr");
const zm = z.math;
const ROBOTO_MONO_TTF = @embedFile("roboto_mono_ttf");
const c = z.colors.Color;
const RT: f32 = 256;
const COLS: usize = 3;
const ROWS: usize = 2;
const State = struct {
font: z.Font,
rt: z.RenderTexture = .{},
frame_count: usize = 0,
};
pub var zimr_app: z.App = .{};
pub fn main() !void {
try zimr_app.run(.{
.window = .{
.title = "zimr - WebGPU - render texture",
.width = 800,
.height = 450,
.scale_mode = .responsive,
},
}, State, initState, update);
}
fn initState(gpa: std.mem.Allocator, f: *z.Frame) !State {
return .{ .font = try z.loadFont(f, gpa, ROBOTO_MONO_TTF, 24) };
}
/// Draw the animated scene into the render texture (RT-pixel coordinates).
fn drawScene(
gl: anytype,
font: z.Font,
t: f32,
) void {
var i: usize = 0;
while (i < 4) : (i += 1) {
const fi: f32 = zm.float(i);
const sz: f32 = 130.0 - fi * 24.0;
const rot: f32 = t * 40.0 + fi * 90.0;
const col: z.colors.Color = z.colorFromHSV(@mod(t * 30.0 + fi * 60.0, 360.0), 0.7, 0.95);
z.drawRectanglePro(
gl,
.{ .x = RT * 0.5, .y = RT * 0.5, .width = sz, .height = sz },
.{ sz * 0.5, sz * 0.5 },
rot,
col.fade(0.7),
);
}
z.drawCircleV(gl, .{ RT * 0.5, RT * 0.5 }, 26.0 + @sin(t * 2.0) * 12.0, c.init(245, 248, 252, 255));
z.drawText(gl, font, "RT", 10, 8, 22, c.init(200, 210, 230, 230));
}
fn update(f: *z.Frame, s: *State) void {
s.frame_count += 1;
if (s.rt.color == .invalid) {
s.rt = z.loadRenderTexture(f.gl, @intFromFloat(RT), @intFromFloat(RT));
}
const t: f32 = f.time.time;
const w: f32 = f.window.widthf();
const h: f32 = f.window.heightf();
z.beginDrawing(f.gl);
z.clearBackground(f.gl, c.init(8, 9, 14, 255));
// 1. Render the animated scene once, into the offscreen render texture.
z.beginTextureMode(f.gl, s.rt, c.init(20, 28, 48, 255));
drawScene(f.gl, s.font, t);
z.endTextureMode(f.gl);
// 2. Composite that single texture to the screen as a tinted grid.
const gap: f32 = 6.0;
const cell_w: f32 = (w - gap * (COLS + 1)) / @as(f32, COLS);
const cell_h: f32 = (h - gap * (ROWS + 1)) / @as(f32, ROWS);
var ry: usize = 0;
while (ry < ROWS) : (ry += 1) {
var cx: usize = 0;
while (cx < COLS) : (cx += 1) {
const idx: f32 = zm.float(ry * COLS + cx);
const dx: f32 = gap + zm.float(cx) * (cell_w + gap);
const dy: f32 = gap + zm.float(ry) * (cell_h + gap);
const tint: z.colors.Color = z.colorFromHSV(@mod(t * 20.0 + idx * 55.0, 360.0), 0.5, 1.0);
z.drawTextureRec(f.gl, s.rt.asTexture(), 0, 0, 1, 1, dx, dy, cell_w, cell_h, tint);
}
}
z.drawText(f.gl, s.font, "render texture: offscreen scene drawn 6x", 10, 10, 14, c.init(210, 214, 224, 230));
z.endDrawing(f.gl);
}