For real-time ray tracing, 1 SPP paired with spatiotemporal denoising (SVGF) offers the most favorable balance between perceptual quality and performance — making higher sampling rates redundant in frame-time-constrained applications. This study evaluates the incremental performance cost and visual impact of ray-traced shadows, reflections, and ambient occlusion in a Vulkan-based hybrid renderer. By analyzing G-buffer-guided ray generation and a 5-level à-trous wavelet SVGF denoiser across both Pre-Lighting and Post-Lighting architectures, we demonstrate that prioritizing feature coverage over sample density yields the best quality-per-millisecond tradeoff on modern hardware.
Hybrid Ray Tracing in Vulkan
Incremental performance and quality analysis of ray-traced shadows and reflections in a custom Vulkan renderer.
Visual Parity
Research Methodology
The project isolates ray tracing overhead by comparing a rasterization baseline (1.35 ms on an AMD RX 9070 XT at 1080p) against incrementally added ray-traced features. The evaluation utilizes two distinct environments: the diffuse-heavy Sponza scene and the specular-heavy Metal-Spheres scene.
- Rasterization Baseline: A traditional deferred pipeline providing primary visibility and evaluating screen-space fallbacks like SSAO and SSR when ray tracing is disabled.
- G-Buffer Guidance vs. Unguided: Rays are primarily seeded from reconstructed world-space positions to reduce BVH traversal overhead, directly compared against unguided object-space dispatching.
- Denoising Architectures: The engine dynamically supports both Pre-Lighting (denoising raw signals before shading) and Post-Lighting (denoising the fully lit signal) SVGF integrations for comparative analysis.
Performance Data — Sponza @ 1080p (AMD RX 9070 XT)
| Configuration | Total GPU (ms) | Overhead |
|---|---|---|
| Raster Only (Baseline) | 1.35 ms | — |
| RT Shadows (1 SPP + SVGF) | 1.88 ms | +0.53 ms |
| RT Reflections (1 SPP + SVGF) | 2.05 ms | +0.70 ms |
| RT Full Pipeline (1 SPP, Post-Lighting) | 2.24 ms | +0.89 ms |
| RT Full Pipeline (8 SPP, Post-Lighting) | 8.80 ms | +7.45 ms |
A complete hybrid pipeline—running ray-traced shadows, reflections, and ambient occlusion simultaneously at 1 SPP—can be achieved for a highly efficient 2.24 ms using the unified Post-Lighting architecture. This Post-Lighting approach cuts the rendering overhead in half compared to Pre-Lighting (4.52 ms) by eliminating secondary intermediate buffers for diffuse-dominated scenes.
Pushing the ray budget to 8 SPP skyrockets the total GPU cost to 8.80 ms. The data shows that increasing sample density yields negligible visual improvements post-denoising, proving that 1 SPP + SVGF is the optimal real-time target.
Implementation Detail — Inline Ray Tracing for RTAO
While G-buffer guidance improves traversal coherence for complex effects like reflections, the evaluation revealed that for short, localized queries like Ambient Occlusion, the fixed architectural overhead of Pipeline Ray Tracing (SBTs and context switching) outweighs the coherence benefits. In these cases, utilizing lightweight Inline Ray Tracing (rayQueryEXT) without G-buffer guidance proved significantly faster.
// GLSL Inline Ray Tracing for short-range occlusion
rayQueryEXT rayQuery;
rayQueryInitializeEXT(rayQuery, topLevelAS, gl_RayFlagsTerminateOnFirstHitEXT, 0xFF, origin, tMin, direction, tMax);
while(rayQueryProceedEXT(rayQuery)) {}
if (rayQueryGetIntersectionTypeEXT(rayQuery, true) != gl_RayQueryCommittedIntersectionNoneEXT) {
// Geometry intersected, pixel is occluded
occlusionFactor = 0.0;
}
Critical Reflection
The most technically demanding part of this project was not the ray tracing itself, but the explicit and unforgiving management of synchronization in Vulkan. Safely transitioning G-Buffer attachments between raster and compute pipelines using VkImageMemoryBarriers proved highly complex, ultimately requiring a data-driven modular render graph to automate hazard tracking and dependency management.
Integrating the SVGF pipeline demonstrated that for real-time applications, the ray tracing step is often the easiest part; the reconstruction is where visual fidelity is truly earned. Realizing that spatiotemporal filtering executes the vast majority of the visual convergence across frames fundamentally changed my approach to performance budgeting, highlighting how computationally inefficient high sample counts are for interactive constraints.
Moving forward, future work will focus on correcting albedo demodulation/remodulation steps to ensure strict energy conservation, and exploring Machine Learning-based reconstruction techniques to evaluate their LPIPS quality improvements against the traditional wavelet-based SVGF filter.