Rendered at 14:49:39 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Pannoniae 3 hours ago [-]
Nice article :) Yeah this is basically a tradeoff between CPU and GPU power. There are different types of culling. There's the basic stuff like backface culling (supported in hardware, don't render triangles facing away from you) and frustum culling (don't render objects which your camera doesn't see). These are used in just about every game.
For occlusion culling it's a bit more tricky because you can either do it on the CPU in broadly two ways, either do low-res raycasting / software rendering like in the article on the CPU and cull based on that. This is an adaptive workload, you can give it more threads or CPU power and it scales for better culling which results in less pixels rendered on the GPU.
You can also use GPU culling but that's more complicated to do and that uses the GPU which creates a catch-22 - you want to use GPU culling to reduce GPU load but integrated GPUs don't cope well with compute shaders and memory bandwidth in general, so doing a culling pass might wipe out any culling gains you might have.
And dedicated GPUs have the raw power and memory bandwidth to just submit everything in your frustum and get most of it depth-rejected.
I wonder if it would be even faster to create a connectivity graph on the CPU, like each chunk knows whether a neighbour is visible and vice versa. On rendering the chunk graph is walked and the visible chunks are submitted, kind of like a primitive garbage collector to determine liveness. The culling would be worse but I presume traversing a fairly small (few thousand elements) list is quite a bit cheaper than rendering the "mipped" occlusion boxes, but do let me know if this is wrong.
keyle 4 hours ago [-]
Wow a tech article on HN. What is happening. Are we back in 2022? /s
fullstackwife 4 hours ago [-]
January 2026, thats light years in AI psychosis time units, thats before harness, and auto hill climbing era
For occlusion culling it's a bit more tricky because you can either do it on the CPU in broadly two ways, either do low-res raycasting / software rendering like in the article on the CPU and cull based on that. This is an adaptive workload, you can give it more threads or CPU power and it scales for better culling which results in less pixels rendered on the GPU.
You can also use GPU culling but that's more complicated to do and that uses the GPU which creates a catch-22 - you want to use GPU culling to reduce GPU load but integrated GPUs don't cope well with compute shaders and memory bandwidth in general, so doing a culling pass might wipe out any culling gains you might have.
And dedicated GPUs have the raw power and memory bandwidth to just submit everything in your frustum and get most of it depth-rejected.
I wonder if it would be even faster to create a connectivity graph on the CPU, like each chunk knows whether a neighbour is visible and vice versa. On rendering the chunk graph is walked and the visible chunks are submitted, kind of like a primitive garbage collector to determine liveness. The culling would be worse but I presume traversing a fairly small (few thousand elements) list is quite a bit cheaper than rendering the "mipped" occlusion boxes, but do let me know if this is wrong.