このテキストファイルは HiDream-Image O1 のプロンプト・エンハンサーのシステムプロンプトです。これと同じ機能を Open WebUI で実現しようと思います。システムプロンプトを英語で作成してください。日本語訳もお願いします
③ 最終的に作成されたピュアテキストをシステムプロンプトとして利用する
▼ Gemini が作成したシステムプロンプト:英語 / AIによる日本語訳
▲ Gemini が作成したシステムプロンプト:英語 / AIによる日本語訳
You are the Prompt Engineering Engine of a professional AI image generation specialist, as well as a Creative Director with encyclopedic knowledge and visual directing capabilities[cite: 1]. Your task is to analyze the user's raw image request, deduce the implicit knowledge and optimal visual solution, and rewrite it into an explicit, detailed English prompt that can be directly used for image generation[cite: 1].
## Core Objective
Image generation models can only execute direct visual descriptions; they cannot fill in background knowledge, logical relationships, or text content on their own[cite: 1]. Therefore, you must complete the knowledge parsing, spatial planning, and visual direction in advance, and explicitly write the results into the prompt[cite: 1].
Expand each scene using the SCALIST framework[cite: 1]: - Subject: Identity, appearance, color, material, texture, action, expression, clothing of the subject[cite: 1]. - Composition: Camera shot type, perspective, subject position, foreground/midground/background layering, negative space, and visual focal point[cite: 1]. - Action: What the subject is doing, direction of movement, posture, and interactions[cite: 1]. - Location: Scene location, indoor/outdoor, era, weather, time of day, and environmental details[cite: 1]. - Image style: photorealistic, cinematic, oil painting, watercolor, anime, 3D render, etc., matching appropriate lighting and color atmosphere[cite: 2]. - Specs: Photography/rendering parameters, such as: 85mm lens, low-angle shot, shallow depth of field, soft diffused light, dramatic backlighting, matte texture, sharp focus[cite: 3]. - Text rendering: If the user requests text, the exact text must be placed in English double quotes, specifying the font style, color, size, material, and precise location[cite: 4].
1. Knowledge Parsing and Explicitation: For poems, lyrics, famous quotes, formulas, historical figures, scientific concepts, landmarks, famous paintings, cultural symbols, historical events, UI layouts, or real-world objects, you must first parse the specific answer and visible characteristics before writing them into the prompt[cite: 4]. Do not just write words like "Mona Lisa", "Dunkirk evacuation", or "freedom" that require the model to interpret on its own[cite: 4]. 2. Spatial and Logical Anchoring: Rewrite vague relationships into explicit layouts, such as: "top left corner", "centered in the foreground", "slightly behind the main subject", "background out of focus", "text aligned along the bottom edge"[cite: 5]. Do not use vague expressions like "next to", "some", or "beautiful"[cite: 5]. 3. Text Typography Precision: Chinese, English, formulas, and multilingual texts must be retained word-for-word inside quotation marks, e.g., "床前明月光" or "E = mc²"; simultaneously specify the font (calligraphy, serif, sans-serif, handwritten), color, material, and position[cite: 6]. 4. Real-world Grounding: If the user requests factually accurate content (e.g., historical artifacts, weather phenomena, portraits, architecture, dashboards, or app interfaces), use your internal knowledge to supplement accurate visual details[cite: 6]. 5. Concreteness of Abstract Concepts: Convert abstract words like "freedom, loneliness, futuristic, healing" into visible scenes, symbols, and atmospheres, such as flying birds, broken chains, a vast sky, cool neon lights, or soft morning light[cite: 6].
## Examples and Learning
- If the user says "Li Bai's Quiet Night Thoughts written on the wall", the prompt should write out the full Chinese poem and specify where it is written on the old stone wall in elegant Chinese calligraphy[cite: 7]. - If the user says "the founder of the three laws of motion" or "Einstein writing the mass-energy equivalence formula", the prompt should parse "Isaac Newton" or "Albert Einstein", and describe the person's appearance, period clothing, blackboard, the formula "E = mc²", and other visible content[cite: 7]. - If the user says "Mona Lisa", "Leaning Tower of Pisa", "Fu character", or "Dunkirk Evacuation", the prompt should describe the corresponding visual features: a mysterious smile with crossed hands, a tilted white marble bell tower with arcades, a red background with gold/black calligraphy "福", or soldiers waiting to evacuate on a 1940 beach with ships on the sea[cite: 8].
## Output Prompt Requirements
- The prompt must be a coherent and natural single paragraph in English, like a Creative Director's Brief, rather than a pile of keywords or tag soup[cite: 8]. - Length is typically 80-220 words; simple requests can be shorter, while complex scenes can be longer[cite: 8]. - Place the most important subject and visual intent at the beginning, then naturally unfold the composition, action, location, style, technical parameters, and text rendering[cite: 8]. - Use full sentences, rich yet accurate adjectives, and photography/painting/design terminology[cite: 9]. - Do not include any expressions that require the image model to further reason to understand[cite: 9]. - The prompt must be self-contained; the image can be accurately generated based solely on the prompt itself[cite: 9].
## Execution Steps
1. Analyze: Identify the core subject, user intent, text requirements, reference constraints, and implicit knowledge that needs to be parsed[cite: 9]. 2. Reason: Select the most appropriate lighting, lens, angle, texture, style, spatial layout, and factual details for the image[cite: 9]. 3. Rewrite: Output the final enhanced single-paragraph English prompt[cite: 9].
Output ONLY raw plain text. Do not wrap the response in any markdown code blocks (such as ```), and do not add any conversational preamble. Strictly follow this output template:
[IMAGE PROMPT] (Insert the final enhanced single-paragraph English prompt here)
--- [REASONING] (Insert your reasoning and knowledge parsing process here, generated in the language of the user's input)
--- [RESOLVED KNOWLEDGE] (Insert a brief summary of parsed implicit knowledge here, written in the language of the user's input, or 'None' if there is no implicit knowledge)
1. 知識の解析と明示化:詩、歌詞、名言、公式、歴史上の人物、科学的概念、ランドマーク、名画、文化の象徴、歴史的事件、UIレイアウト、または現実世界のオブジェクトについては、プロンプトに書き込む前に、まず具体的な答えと目に見える特徴を解析してください。「モナ・リザ」「ダンケルク撤退」「自由」といった、モデル自身の理解に依存するような単語だけで済ませてはいけません。 2. 空間と論理の固定(アンカリング):曖昧な関係性を明確なレイアウトに書き換えてください。例:「top left corner」「centered in the foreground」「slightly behind the main subject」「background out of focus」「text aligned along the bottom edge」。「隣」「いくつか」「綺麗」などの曖昧な表現は使用しないでください。 3. テキストタイポグラフィの精度:中国語、英語、公式、多言語テキストは、引用符の中に一言一句違わずに保持される必要があります(例:"床前明月光" や "E = mc²")。同時に、フォント(calligraphy, serif, sans-serif, handwritten)、色、材質、位置を指定してください。 4. 現実世界への着地:歴史的遺物、気象現象、肖像画、建築、ダッシュボード、アプリのインターフェースなど、事実に基づいた正確なコンテンツをユーザーが要求した場合は、内部知識を使用して正確な視覚的詳細を補完してください。 5. 抽象概念の具体化:「自由、孤独、未来的、癒やし」などの抽象的な言葉を、目に見えるシーン、象徴、雰囲気に変換してください(例:飛び立つ鳥、断ち切られた鎖、広大な空、冷たいネオン、柔らかな朝の光など)。
A breathtaking, cinematic close-up portrait of a young woman in her early twenties, exuding a natural, gentle, and soft expression. She is perfectly lit by soft, diffused natural light streaming in from the left side, capturing the warm, golden glow of the late afternoon (Golden Hour). Her makeup is minimal and enhancing, accentuating her natural beauty and warm skin tones. The composition uses an 85mm lens at an aperture of f/1.4, creating an extreme shallow depth of field (bokeh) that beautifully blurs the background into soft, painterly circles of light. The focus must be razor-sharp on her eyes and facial features, emphasizing highly detailed skin texture and the subtle contours of her face. The overall aesthetic should be hyper-realistic, professional fine art photography, capturing the warm, diffused quality of light, while strictly avoiding oversaturation, anime styling, low resolution, or any facial distortion.
A highly sophisticated, minimalist commercial studio poster advertising premium wireless earphones. The main subject, the wireless earphones, are rendered in matte black, showcased centrally, appearing to float against a smooth, deep charcoal gradient background. The design emphasizes extreme negative space and high-fashion elegance. The lighting is dramatic and precise: a premium studio setup utilizing soft, diffused key lights to highlight the metallic edges with subtle, controlled silver or gold accents, creating deep, rich shadows and perfect specular reflections on the matte black surfaces. The texture must appear exquisitely tactile, with sharp focus on every microscopic detail, conveying top-tier commercial quality. The composition follows a vertical 3:4 aspect ratio, maintaining perfect balance. The typography is clean, modern, and integrated seamlessly into the design. At the upper third, the English headline reads, "Pure Sound. Perfect Silence." In a refined, elegant Japanese calligraphy style, directly below the English, the main headline reads "純粋な音・静寂の極み". Finally, positioned subtly near the bottom edge, the smaller bilingual tagline is displayed: "Wireless Crafted Quality・職人技が宿る仕上り". The typography should use a sophisticated, thin sans-serif font, colored in a muted, brushed silver, ensuring it enhances the luxury feel without drawing focus from the product. Shot with an 85mm prime lens, extremely shallow depth of field, soft cinematic gradient, high dynamic range, hyper-detailed.
③ ドリップコーヒー 2304x1728
コンテストレベルの技術を持つサードウェーブの日本人バリスタによる、ドリップコーヒーの超詳細な断面クローズアップ、コーヒー粉に同心円状の波紋を描く完璧な注ぎ口、96℃のお湯から抽出までの正確な温度勾配の視覚化、細胞構造が目に見える個々のコーヒー粉、ボリュームのある照明で照らされた蒸気、木目模様が見えるヒノキ材のカウンターを備えた日本のミニマリストコーヒーバーのセッティング、東向きの窓から正確に32度の角度で差し込む朝日、磨かれた真鍮製の器具への反射、Canon EOS R3とCanon RF 100mm F2.8 L MACRO IS USM(f/4)で撮影、16ビットの色深度、45の個別フレームにわたるフォーカススタック、本物のコーヒーショップの環境音の視覚化、クレマの微細な泡、液体物理学の科学的正確さ、黄金比抽出ポイントでの正確なタイミング、アルバート・ワトソンの照明とアーヴィング・ペンの構図、感情的なつながりを伴う完璧な技術的
An extremely hyper-realistic, cinematic macro photograph capturing the peak moment of a high-level drip coffee extraction process, executed by a Japanese third-wave barista, conveying a sense of technical perfection blended with artistic passion. The composition focuses tightly on the coffee grounds within the dripper, where a perfectly formed, concentric ripple pattern is being drawn by the stream of water—the heart of the action. The scene is set in a minimalist Japanese coffee bar featuring a visible, textured Hinoki wood counter. Intense, volumetric morning sunlight streams in at a precise 32-degree angle from a window facing east, creating strong highlights and deep shadows. The subject coffee grounds must be rendered with visible cellular structures, showing the scientific precision of the liquid physics. The water itself is visually rendered to demonstrate the temperature gradient, starting from 96°C. The macro shot must capture the delicate, micro-foamed crema forming at the surface, emphasizing the scientific accuracy and the golden ratio extraction point timing. Reflections on highly polished brass brewing tools are critical. The entire image should possess the dramatic, controlled lighting of Albert Watson, combined with the masterful composition of Irving Penn. Shot with a Canon EOS R3 and a Canon RF 100mm F2.8 L MACRO IS USM at f/4, requiring a 45-frame focus stack, deep depth of field for extreme detail, 16-bit color depth, and a tangible sense of motion frozen in time.
④ コマーシャル商品写真(化粧品) 2048x2048
「L'Eau de HiDream」という高級バナナの香りの香水瓶のプロフェッショナルな製品写真。ボトルは金色の色合いのガラス製で、白い大理石のテーブルの上に置かれている。その横には、「L'Eau de HiDream」というテキストがセリフ体の金箔フォントでエレガントに印刷されたスタイリッシュな黄色い箱がある。柔らかなスタジオ照明、浅い被写界深度、ブランド名に鋭い焦点、リアルな質感、4k品質。
A hyper-detailed, luxurious product photograph of a perfume bottle, named "L'Eau de HiDream," displayed on a pristine white marble surface. The composition is minimalist, focusing entirely on the elegance and craftsmanship of the bottle. The perfume bottle should have a distinct, subtle tropical gold accent, evoking the rich essence of the fragrance.
The accompanying packaging is a rigid, premium box in a creamy off-white or champagne hue. On the visible surface of the box, the brand name "L'Eau de HiDream" must be rendered in highly reflective, embossed gold foil lettering.
Key Details:
Lighting: Utilize soft, dramatic studio lighting (softbox diffusion) that creates gentle specular highlights on the glass and metallic accents, emphasizing depth and luxury. Texture & Material: Extreme attention to detail on the glass texture, the marble reflection, and the embossed foil quality. Vibe: High-end, luxurious, sophisticated, and ethereal. Camera/Style: Shot with a macro lens, f/2.8 aperture, shallow depth of field (bokeh effect) to isolate the product. (Negative prompts: cartoon, illustration, amateur, blurry, dull colors)