A couple of things I've found that work:
1) Don't use the prompt improver, it overrides your settings way too often. Instead refer to the H3 guide and hand write it, shorter prompts actually work much better than overly detailed prompts.
2) Break you video into 7 second generations, don't try to go too long. It's not just faster to render and fix if you get messed up output, but a basic rule of cinematography is to change camera shots about every 7 seconds. It keeps the audience better engaged. This means 99% of your camera work will be set to "Static Shot" in your prompt. You can cut to different angles, but camera movement is for more dramatic scenes, filler scenes and establishing shots. That will keep the character's face in frame which will prevent drift.
3) Create a dozen or more "Character" profiles for each character. The paper-doll method, you have your base character image in high definition, then you create new profiles for different wardrobe, etc. So let's say your character is "Jim", and for this scene he ill be shot in profile, wearing a red shirt, black pants, white sneakers and sunglasses. You have a character profile for that exact arrangement, with a still image of him standing in profile in that exact attire. You'll have a different one of him in that same attire for a close-up headshot, another one for a distant shot with the full body in frame, etc. So one "Character" in the UI will be dozens of different templates created for the shot you're trying to get.
4) Use Image Generation for your first frame for every ~7 second shot, and story-board it. Plot out a timeline in a spreadsheet, each row in 7 second intervals that contains the prompt, a column for camera position, lighting, setting, character definition, ambient sounds, etc. Structuure it to generate the H3 markup per teh schema (found on HuggingFace). If you use Excel for storyboarding,you can use the data validation function to get a nice drop-down list for a lot of default options liek camera positions, so you always have the correct syntax.
5) Record your own voice for any dialogue, even if that's not what you use in the final cut. What you want is the time it takes to say the line. If you plug 45 seconds of dialogue into a 7 second clip, it's not going to work. Likewise, if you say 1 second of dialogue, it will often fill in with gibberish. I f possible pre-record the audio and generate the video with an audio reference to keep things synced.
6) H3 is insenely good at facial expressions, but that's one area you want to be very descriptive. Not just "Happy", "Sad", "Angry". Maybe watch a reference video and try to describe it completely... "Jim lets out and abrupt, muffled chuckle to disguise his trepidation".
7) A lot of the LORAs work against you. They're best used for single shots as applicable, none should be used by default, and you call on them only as needed. If you need a close-up of hands, then you use a LORA for hands, only for that shot.
8) For longer 30+ second shots, still break them up into 7-10 second segment, and use the last frame of the first generation as the first frame of the second.
9) Do NOT queue up a prompt with 10+ windows, it will drift rapidly and take forever to generate. Each one shot is it's own prompt, run one at a time. Every 2-3 shots, I kill the app and restart, as it seems to get bogged down after a few windows are run. Not sure if there's a memory leak or if it's just hanging on to too much context from previous generations, but it will be loads faster after a fresh restart. As long as you have a storyboard, this isn't going to slow you down, quite the opposite.
10) Face refinement works, but keep it to about 60%. Try to limit the number of characters in frame at one time. All character should be named and have a reference profile. When you have more than 3-4 characters on screen at once, use image gen for a reference photo and try to keep some of them in the backdrop out of camera focus, they can be seen there, but not in the detail as the focus characters. You can define the camera's focal length so those in the background naturally look a little more blurry, and their faces don't distort because it's not trying to add greater detail there.
11) Your character's face is either in frame for that window or it's not, don't occulde their face and then have them turn toward the camera again, or it will be distorted. You can "hack" this if you need them to look dramatically into the camera by actually having them look away instead, then reversing that segment in editing later. But the second their face is out of frame, the reference is lost and when the look at the camera again, it will be slightlu distorted. This also happens a lot with body types and clothing details as well. Once something is occluded from frame, don't bring it back until the next window, which is why shorter window lengths help tremendously. Then you're always starting from a clean reference image that has all it's detail in tact.
12) The random all-nighter fishing expidition works. Queue up 20+ frames at 30 seconds, make sure all reference the source images so they don't drift. Then just run a long prompt with various actions, movements, facical expressions, etc. The otput will be garbage, but you can play it in VLC and screen cap usable frames for later reference. You'll have an hour of crap video, but a hundred good still shots you can reference later to make better image generations or even use directly in more deliberate video scenes. You'd be suprised some fo the awesome shots you randomly get but would be difficult to prompt at all without the screncap as a reference.
So baically, do 10x more prep work on the video upfront before clicking the generate button. That investment in time will pay dividends later on when you're not re-running lots of generations to deal with minor flaws. Hope that helps.
