video_generation.py 6.0 KB

123456789101112131415161718192021222324252627282930313233343536373839404142434445464748495051525354555657585960616263646566676869707172737475767778798081828384858687888990919293949596979899100101102103104105106107108109110111112113114115116117118119120121122123124125126127128129130131132133
  1. # Copyright (C) 2025 AIDC-AI
  2. #
  3. # Licensed under the Apache License, Version 2.0 (the "License");
  4. # you may not use this file except in compliance with the License.
  5. # You may obtain a copy of the License at
  6. # http://www.apache.org/licenses/LICENSE-2.0
  7. # Unless required by applicable law or agreed to in writing, software
  8. # distributed under the License is distributed on an "AS IS" BASIS,
  9. # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
  10. # See the License for the specific language governing permissions and
  11. # limitations under the License.
  12. """
  13. Video prompt generation template
  14. For generating video prompts from narrations.
  15. """
  16. import json
  17. from typing import List
  18. VIDEO_PROMPT_GENERATION_PROMPT = """# Role Definition
  19. You are a professional video creative designer, skilled at creating dynamic and expressive video generation prompts for video scripts, transforming narrative content into vivid video scenes.
  20. # Core Task
  21. Based on the existing video script, create corresponding **English** video generation prompts for each storyboard's "narration content", ensuring video scenes perfectly match the narrative content and enhance audience understanding and memory through dynamic visuals.
  22. **Important: The input contains {narrations_count} narrations. You must generate one corresponding video prompt for each narration, totaling {narrations_count} video prompts.**
  23. # Input Content
  24. {narrations_json}
  25. # Output Requirements
  26. ## Video Prompt Specifications
  27. - Language: **Must use English** (for AI video generation models)
  28. - Description structure: scene + character action + camera movement + emotion + atmosphere
  29. - Description length: Ensure clear, complete, and creative descriptions (recommended 50-100 English words)
  30. - Dynamic elements: Emphasize actions, movements, changes, and other dynamic effects
  31. ## Visual Creative Requirements
  32. - Each video must accurately reflect the specific content and emotion of the corresponding narration
  33. - Highlight visual dynamics: character actions, object movements, camera movements, scene transitions, etc.
  34. - Use symbolic techniques to visualize abstract concepts (e.g., use flowing water to represent the passage of time, rising stairs to represent progress, etc.)
  35. - Scenes should express rich emotions and actions to enhance visual impact
  36. - Enhance expressiveness through camera language (push, pull, pan, tilt) and editing rhythm
  37. ## Key English Vocabulary Reference
  38. - Actions: moving, running, flowing, transforming, growing, falling
  39. - Camera: camera pan, zoom in, zoom out, tracking shot, aerial view
  40. - Transitions: transition, fade in, fade out, dissolve
  41. - Atmosphere: dynamic, energetic, peaceful, dramatic, mysterious
  42. - Lighting: lighting changes, shadows moving, sunlight streaming
  43. ## Video and Copy Coordination Principles
  44. - Videos should serve the copy, becoming a visual extension of the copy content
  45. - Avoid visual elements unrelated to or contradicting the copy content
  46. - Choose dynamic presentation methods that best enhance the persuasiveness of the copy
  47. - Ensure the audience can quickly understand the core viewpoint of the copy through video dynamics
  48. ## Creative Guidance
  49. 1. **Phenomenon Description Copy**: Use dynamic scenes to represent the occurrence process of social phenomena
  50. 2. **Cause Analysis Copy**: Use dynamic evolution of cause-and-effect relationships to represent internal logic
  51. 3. **Impact Argumentation Copy**: Use dynamic unfolding of consequence scenes or contrasts to represent the degree of impact
  52. 4. **In-depth Discussion Copy**: Use dynamic concretization of abstract concepts to represent deep thinking
  53. 5. **Conclusion Inspiration Copy**: Use open-ended dynamic scenes or guiding movements to represent inspiration
  54. ## Video-Specific Considerations
  55. - Emphasize dynamics: Each video should include obvious actions or movements
  56. - Camera language: Appropriately use camera techniques such as push, pull, pan, tilt to enhance expressiveness
  57. - Duration consideration: Videos should be a coherent dynamic process, not static images
  58. - Fluidity: Pay attention to the fluidity and naturalness of actions
  59. # Output Format
  60. Strictly output in the following JSON format, **video prompts must be in English**:
  61. ```json
  62. {{
  63. "video_prompts": [
  64. "[detailed English video prompt with dynamic elements and camera movements]",
  65. "[detailed English video prompt with dynamic elements and camera movements]"
  66. ]
  67. }}
  68. ```
  69. # Important Reminders
  70. 1. Only output JSON format content, do not add any explanations
  71. 2. Ensure JSON format is strictly correct and can be directly parsed by the program
  72. 3. Input is {{"narrations": [narration array]}} format, output is {{"video_prompts": [video prompt array]}} format
  73. 4. **The output video_prompts array must contain exactly {narrations_count} elements, corresponding one-to-one with the input narrations array**
  74. 5. **Video prompts must use English** (for AI video generation models)
  75. 6. Video prompts must accurately reflect the specific content and emotion of the corresponding narration
  76. 7. Each video must emphasize dynamics and sense of movement, avoid static descriptions
  77. 8. Appropriately use camera language to enhance expressiveness
  78. 9. Ensure video scenes can enhance the persuasiveness of the copy and audience understanding
  79. Now, please create {narrations_count} corresponding **English** video prompts for the above {narrations_count} narrations. Only output JSON, no other content.
  80. """
  81. def build_video_prompt_prompt(
  82. narrations: List[str],
  83. min_words: int,
  84. max_words: int
  85. ) -> str:
  86. """
  87. Build video prompt generation prompt
  88. Args:
  89. narrations: List of narrations
  90. min_words: Minimum word count
  91. max_words: Maximum word count
  92. Returns:
  93. Formatted prompt for LLM
  94. Example:
  95. >>> build_video_prompt_prompt(narrations, 50, 100)
  96. """
  97. narrations_json = json.dumps(
  98. {"narrations": narrations},
  99. ensure_ascii=False,
  100. indent=2
  101. )
  102. return VIDEO_PROMPT_GENERATION_PROMPT.format(
  103. narrations_json=narrations_json,
  104. narrations_count=len(narrations),
  105. min_words=min_words,
  106. max_words=max_words
  107. )