Phyrex
Phyrex|Aug 25, 2026 16:51
There are still many experts on X, and under their guidance, I have switched from RTX PRO 6000 Blackwell Max-Q 96GB X4 to the standard version. Although the price of Max-Q and the standard version is the same due to the current increase in memory, the standard version is much more difficult to configure. It is only okay to run dual SIM cards, but it is difficult to handle running four SIM cards. In addition, I carefully studied "Huangguo" and roughly reverse its approach. I used an 8-minute video to break it down, and without including the opening and ending credits, there are about 117 shots in total. The median length of the shots is about 3.2 seconds, with about 72% of the shots not exceeding 5 seconds, about 92% of the shots not exceeding 8 seconds, and about 95% of the shots not exceeding 10 seconds. By analyzing the traces of 24fps to 30fps conversion, it is found that some shots will have a repeated frame with a height close to the previous frame every 5 frames on average, which is consistent with the feature of 24fps material being converted to 30fps through repeated frames. Moreover, the repeated frame positions of different shots are not exactly the same, indicating that these shots are likely to be generated and processed separately before being finally placed on the 30fps editing timeline. This indicates that Wan 2.2 is likely to be used at the bottom layer, but the official output of Wan 2.2 is 24fps. Combined with lens length, character drift, clothing changes, scene structure changes, and image generated video features, Wan 2.1, Wan 2.2, or models based on Wan fine-tuning can be used, so Wan is sufficient. Huangguo most likely does not have a model that can directly generate 8-minute adult videos. The production method should first use a big language model to write the script and split the shots, then fix the characters with LoRA or reference images, use image models to generate high-quality keyframes one by one, and then hand them over to I2V models like Wan to generate videos of 3 to 8 seconds. Dialogue shots are separately driven by mouth movements and facial expressions, while adult shots use specially tuned models or LoRA. Complex two person movements may also include reference videos, poses, or video to video. Finally, perform super-resolution, flicker removal, 24fps to 30fps conversion, dubbing, subtitles, system interface, and editing uniformly. The adult shots are likely not generated using Wan, or in other words, pure adult shots are not generated using Wan, because the single frame quality of the adult part is very high, and the long-term action quality is average. When capturing a single frame, the skin, lighting, body proportions, and composition of the character can be very fine. But when playing continuously, you will see many bugs. I suspect that adult shots are likely to undergo three layers of processing. Firstly, the keyframes of adult images are processed using image checkpoints or LoRA trained on adult data, resulting in relatively accurate static images of poses, characters, and scenes. Then there is adult video fine-tuning, using Wan series or similar video models such as adult LoRA, low rank fine-tuning, or full-scale fine-tuning, to enable the model to understand weaker adult actions in the original weights. Finally, there is action reference. Some interactions between two or more people may involve inputting action reference videos, posture sequences, or depth information. The image is generated, but the motion trajectory may come from real footage or other videos. So the most likely complete production line is: LLM Writing Plot 1. Automatically split shots and lines 2. Establish a male and female role setting diagram 3. Train LoRA or establish a reference library for the main character 4. Establish scene templates 5. The image model generates keyframes shot by shot 6. Manual screening, hand repair, face repair, and text repair 7. Use Wan I2V or TI2V for regular lenses 8. Use Wan S2V or post production Lip sync for dialogue shots 9. Adult lenses use adult to fine tune towards Wan or LoRA 10. Partial complex action input for real human motion reference 11. Hypergrading, debounce, and 24 → 30fps 12. TTS、 Music, sound effects, and mixing 13. Subtitles UI、 Advertisements and watermarks 14. Cut into a complete 8-minute short play Welcome everyone to discuss. @Gate Crypto、 US stocks, Hong Kong stocks, South Korean stocks, gold CFD、 Predicting one-stop trading in the market
+2
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads