Training with an automated data pipeline
Voyager is based on Tencent's early Hunyuanworld 1.0, released in July. Voyager is also part of Tencent's broader “Hunyuan” ecosystem. This includes the 3D generation Hunyuan3D-2 model from text and previously treated for video integration.
To train Voyagers, researchers have developed software that automatically analyzes existing videos to process camera movements and calculates the depth of all frames. The system processed over 100,000 video clips from both the actual recordings and the aforementioned unrealistic engine rendering.
Illustration of Voyager World Creation Pipeline.
Credit: Tencent
The model needs to run serious computing power and requires at least 60GB of GPU memory at 540p resolution, but Tencent recommends 80GB for better results. Tencent exposed the weights of the face-hugging model and included code that works with both single and multi-GPU setups.
This model comes with notable licensing restrictions. Like Tencent's other Hunyuan models, the license bans its use in the European Union, the UK and South Korea. Additionally, commercial deployments that serve more than 100 million active users each month require a separate license from Tencent.
In a world score benchmark developed by researchers at Stanford University, Voyager reportedly achieved a top overall score of 77.62 compared to 72.69 for Wonderworld and 62.15 for Cogvideox-I2V. The model reportedly excels in object control (66.92), style consistency (84.89), and subjective quality (71.09), but came in second in camera control (85.95) behind Wonderworld's 92.98. WorldScore evaluates world-generation approaches across multiple criteria, including 3D consistency and content adjustment.
While these self-reported benchmark results appear to be promising, the wider developments still face challenges due to the computational muscles involved. For developers who need faster processing, the system supports parallel inference across multiple GPUs using the XDIT framework. Running on 8 GPUs provides 6.69 times faster processing speed than a single GPU setup.
Given the processing power required and the limitations of generating a long and consistent “world”, it may take some time to see a real-time interactive experience using similar techniques. However, as we have seen in experiments such as Google's Genie, we are potentially witnessing very early stages into new interactive, generative art forms.
