概要
よい研究も,その魅力が伝わらなければ多くの人には届かない.論文やコードとあわせて公開される「Project Page」は,研究成果を広く届けるための重要な入口である.一方で,どのようなページが研究の認知や理解につながるのかは,十分に整理されていない.
本プロジェクトでは,CVPR 2023–2025の論文を対象に,Project Pageの公開と引用数の関係,AIによる自動評価と人間の主観評価の関係,そして「わかりやすさ」を左右するデザイン要因を調査した.その結果,ページを公開している論文は引用数が高い傾向にあり,単に情報を多く載せることよりも,動画や図,レイアウト,図と文章のバランスを工夫することが重要だと分かった.
引用数
Project Pageを公開すると,論文はより引用されるのか?
Project Pageの公開が研究の認知に貢献するかを検証するため,引用数を尺度としてProject Pageの有無による差を比較した. CVPRの3年間の論文を対象に各論文の引用数とProject Pageの有無を収集し,各年度の引用数の中央値および四分位範囲を可視化した. 3年分の結果によりProject Pageの公開が引用数向上に貢献する結果が得られた.
定量評価
AIによる自動評価は,人の「わかりやすい」と一致するか?
Paper2Web は,Project Pageの品質を,デザイン性(Aesthetic)・情報の完全性(Completeness)・リンク品質(Connectivity)・インタラクティブ性(Interactive)といった観点から自動評価する枠組みと,論文の内容がProject Pageからどれだけ伝わるかをクイズ形式で測る PaperQuiz を提案している. 本調査ではこれらの既存指標を実際に公開されているProject Pageに適用し,人間の主観的な「わかりやすさ」をどの程度捉えられるかを定量的に検証した.
自動指標と主観評価の相関
指標名をクリックすると評価方法を表示
値は相関係数の平均 ± 標準偏差.右に伸びるほど,人間による高評価と同じ傾向を示す.
まとめ
「たくさん書く」より,「見やすく,簡潔に伝える」. 文字による情報量が多いProject Pageよりも,デザイン性を重視した簡潔なProject Pageのほうが高く評価される傾向が見られた. 実際のProject Pageを用いた比較は定性評価を参照.
評価したProject Pageを見る
54 pages
サムネイルを選ぶと,そのProject Pageの各指標の点数と人手評価の順位,そして実際のページを確認できる.
デザイン性
–
情報の完全性
–
リンク品質
–
インタラクティブ性
–
情報量
–
人間の順位
–
下の枠には実際のProject Pageを埋め込んでいる.枠内でそのまま操作できる.
このページは枠内に表示できません
実際のページを開く ↗Paper Quizで生成される質問の例 ▼
BARD-GS (CVPR 2025) のプロジェクトページで生成された質問.
Which institution is Robin Rombach affiliated with?
Robin Rombachはどの機関に所属していますか?
A. Stanford University
A. スタンフォード大学
B. LMU Munich
B. LMUミュンヘン
C. University of California, Berkeley
C. カリフォルニア大学バークレー校
D. Massachusetts Institute of Technology
D. マサチューセッツ工科大学
During internships at which company did Andreas, Robin, and Tim perform the work described in the paper?
AndreasとRobin、Timは本論文の研究をどの企業のインターン中に行いましたか?
A. Apple Inc.
A. Apple Inc.
B. Meta Platforms
B. Meta Platforms
C. NVIDIA
C. NVIDIA
D. Google AI
D. Google AI
What is the primary reason video modeling has lagged behind image modeling despite breakthroughs in generative models?
生成モデルの進歩にも関わらず、動画モデリングが画像モデリングに遅れている主な理由は何ですか?
A. Significant computational cost and lack of large-scale, general video datasets.
A. 大幅な計算コストと大規模な汎用動画データセットの不足
B. Insufficient computing power for video data.
B. 動画データに対する計算能力の不足
C. Lack of sophisticated algorithms for video processing.
C. 動画処理のための高度なアルゴリズムの欠如
D. Limited interest in video generation applications.
D. 動画生成アプリへの関心の低さ
One of the real-world problems addressed by the Video LDM is creative content creation using what method?
Video LDMが取り組む現実の課題の一つとして、どの手法によるコンテンツ制作が挙げられますか?
A. Manual frame-by-frame animation
A. フレームごとの手動アニメーション
B. 3D rendering engines
B. 3Dレンダリングエンジン
C. Style-transfer algorithms
C. スタイル転送アルゴリズム
D. Text-to-video modeling
D. テキストから動画へのモデリング
What is a key insight for efficiently training a video generation model, as presented in the paper?
動画生成モデルを効率的に学習するための重要な洞察として論文で紹介されているものは何ですか?
A. Training a diffusion model directly in pixel space from scratch.
A. 拡散モデルをピクセル空間でゼロから直接学習する
B. Developing entirely new diffusion model architectures for video.
B. 動画用に全く新しい拡散モデルアーキテクチャを開発する
C. Relying solely on autoregressive transformers for video generation.
C. 動画生成に自己回帰トランスフォーマーのみを使用する
D. Re-using a pre-trained, fixed image generation model.
D. 事前学習済みの固定画像生成モデルを再利用する
How do the authors propose to transform an LDM image generator into a video generator?
著者らはLDM画像生成器を動画生成器に変換するためにどのような方法を提案していますか?
A. By using larger image datasets for pre-training.
A. 大規模な画像データセットを事前学習に使用する
B. By introducing a temporal dimension to the latent space diffusion model and fine-tuning on encoded image sequences.
B. 潜在空間拡散モデルに時間次元を導入し符号化画像シーケンスでFTする
C. By increasing the batch size of image training data.
C. 画像学習データのバッチサイズを増やす
D. By re-training the entire LDM architecture on video data.
D. LDMアーキテクチャ全体を動画データで再学習する
What is the resolution and frame rate of the in-house dataset of real driving scene (RDS) videos used in this work?
本研究で使用されたリアル走行シーン(RDS)動画の社内データセットの解像度とフレームレートは?
A. 1920×1080 at 60 fps
A. 1920×1080、60fps
B. 512×1024 at up to 30 fps
B. 512×1024、最大30fps
C. 1024×512 at 15 fps
C. 1024×512、15fps
D. 720×1280 at 24 fps
D. 720×1280、24fps
What is the total number of videos in the Real Driving Scene (RDS) dataset mentioned?
リアル走行シーン(RDS)データセットの総動画数は何本ですか?
A. 100,000 videos
A. 10万本
B. 683,060 videos
B. 683,060本
C. 50,000 videos
C. 5万本
D. 1,000,000 videos
D. 100万本
What component of the image LDM remains fixed while only the parameters of the temporal layers are optimized during fine-tuning?
ファインチューニング中に時間層のパラメータのみ最適化される際、画像LDMのどのコンポーネントが固定されたままですか?
A. The spatial layers of the image backbone
A. 画像バックボーンの空間層
B. The latent space dimensions
B. 潜在空間の次元
C. The noise schedule parameters
C. ノイズスケジュールパラメータ
D. The entire decoder module
D. デコーダモジュール全体
What are the two different kinds of temporal mixing layers implemented in the Video LDM?
Video LDMに実装されている2種類の時間混合層とは何ですか?
A. Spatial convolution and recurrent networks
A. 空間畳み込みと再帰型ネットワーク
B. Temporal attention and residual blocks based on 3D convolutions
B. 時間的アテンションと3D畳み込みに基づく残差ブロック
C. Gated recurrent units and long short-term memory
C. ゲート付き再帰ユニットとLSTM
D. Fully connected layers and self-attention
D. 全結合層と自己アテンション
What is the guidance scale used for the Text-to-Video LDM keyframes model during sampling?
テキストから動画へのLDMキーフレームモデルのサンプリング時のガイダンススケールは何ですか?
A. 8.0
A. 8.0
B. 5.0
B. 5.0
C. 1.0
C. 1.0
D. 2.0
D. 2.0
What is the number of diffusion steps used for the Diffusion Setup in the Driving (Video) LDM?
Driving Video LDMの拡散セットアップで使用される拡散ステップ数は何ですか?
A. 500
A. 500
B. 250
B. 250
C. 1000
C. 1000
D. 2000
D. 2000
What metric is used for frame-wise evaluation of the models?
モデルのフレームごとの評価に使用される指標は何ですか?
A. Peak Signal-to-Noise Ratio (PSNR)
A. ピーク信号対雑音比(PSNR)
B. Fréchet Inception Distance (FID)
B. フレシェ開始距離(FID)
C. Mean Squared Error (MSE)
C. 平均二乗誤差(MSE)
D. Structural Similarity Index (SSIM)
D. 構造的類似度指標(SSIM)
According to the text, for what specific aspect is human evaluation used, apart from quantitative metrics?
定量的指標以外に人間評価が使用される特定の側面は何ですか?
A. Computational efficiency
A. 計算効率
B. Model complexity
B. モデルの複雑さ
C. Training convergence speed
C. 学習の収束速度
D. Realism of generated videos
D. 生成動画のリアリティ
What is the FVD score obtained by 'Ours (cond.)' when compared with LVG on RDS, as shown in Table 1?
表1に示されるRDS上でのLVGとの比較における「Ours (cond.)」のFVDスコアは何ですか?
A. 356
A. 356
B. 478
B. 478
C. 389
C. 389
D. 534.17
D. 534.17
For the 'Video discriminator only' method in decoder fine-tuning, what is the Reconstruction FVD score?
デコーダFTの「Video discriminator only」手法の再構成FVDスコアはいくつですか?
A. 9.04
A. 9.04
B. 32.94
B. 32.94
C. 51.01
C. 51.01
D. 9.17
D. 9.17
What does Figure 2, 'Temporal Video Fine-Tuning', primarily illustrate?
図2「Temporal Video Fine-Tuning」が主に示しているのは何ですか?
A. How pre-trained image diffusion models are turned into temporally consistent video generators.
A. 事前学習済み画像拡散モデルが時間的に一貫した動画生成器に変換される方法
B. The impact of hyperparameter tuning on video quality.
B. ハイパーパラメータチューニングが動画品質に与える影響
C. The process of generating key frames for long videos.
C. 長動画のキーフレーム生成プロセス
D. The architecture of the spatial layers in the LDM.
D. LDMの空間層のアーキテクチャ
What does Figure 1, the video LDM samples, demonstrate regarding the model's capabilities?
図1のVideo LDMサンプルはモデルの何の能力を示していますか?
A. Both text-to-video generation and real driving scene video generation.
A. テキストから動画への生成とリアル走行シーン動画生成の両方
B. The computational cost savings of the LDM paradigm.
B. LDMパラダイムによる計算コスト削減
C. Only real driving scene video generation.
C. リアル走行シーン動画生成のみ
D. Only text-to-video generation.
D. テキストから動画への生成のみ
Which ablation study result indicates a significant degradation in both FID and FVD when compared to the Video LDM?
FIDとFVDの両方でVideo LDMより大幅な劣化を示すアブレーション研究の結果はどれですか?
A. End-to-End LDM
A. End-to-End LDM
B. Ours (context-guided)
B. Ours (context-guided)
C. Attention-only
C. Attention-only
D. Pixel-baseline
D. Pixel-baseline
What key finding resulted from comparing video fine-tuned pixel-space upsampler with independent frame-wise image upsampling?
動画FTされたアップサンプラーと独立したフレームごとのアップサンプリングを比較した重要な発見は何ですか?
A. Pixel-space upsampling is always superior to latent space upsampling.
A. ピクセル空間アップサンプリングは常に潜在空間より優れている
B. Temporal alignment of the upsampler is crucial for high performance.
B. アップサンプラーの時間的整合性は高いパフォーマンスに不可欠
C. FID was significantly affected by temporal alignment.
C. FIDは時間的整合性によって大きく影響された
D. Independent upsampling significantly improved FVD.
D. 独立したアップサンプリングはFVDを大幅に改善した
What is one of the key design choices of the presented Video Latent Diffusion Models?
提示されたVideo Latent Diffusion Modelsの主要な設計選択の一つは何ですか?
A. To use only pixel-space diffusion models for all video generation tasks.
A. 全動画生成タスクにピクセル空間拡散モデルのみを使用する
B. To solely rely on new, untried diffusion model architectures.
B. 新しい未試行の拡散モデルアーキテクチャのみに依存する
C. To build on pre-trained image diffusion models and turn them into video generators by temporally video fine-tuning them.
C. 事前学習済み画像拡散モデルを基に時間的動画FTで動画生成器に変換する
D. To forego high-resolution video generation in favor of computational efficiency.
D. 計算効率のために高解像度動画生成を断念する
Apart from high-resolution video generation, what specific application is mentioned as a benefit of Video LDM, particularly in the context of autonomous driving?
高解像度動画生成の他に、自律走行の文脈でVideo LDMの利点として言及されている応用は何ですか?
A. Traffic flow prediction
A. 交通流予測
B. Autonomous vehicle control
B. 自動運転車制御
C. Real-time object detection
C. リアルタイム物体検出
D. Serving as simulators in autonomous driving research
D. 自律走行研究のシミュレータとして機能する
What is a potential negative ethical implication highlighted regarding enhanced versions of the model achieving higher quality?
モデルの強化版が高品質を達成することで生じる潜在的な否定的倫理的含意は何ですか?
A. Difficulty in deployment on mobile devices
A. モバイルデバイスへの展開の難しさ
B. Increased computational costs
B. 計算コストの増大
C. Potential for generating deceptively real videos for malicious purposes
C. 悪意ある目的のために欺瞞的にリアルな動画を生成する可能性
D. Limited artistic expression
D. 芸術的表現の制限
An important direction for future work mentioned is training large-scale generative models with what kind of data?
将来の研究の方向として、どのようなデータで大規模生成モデルを学習させることが重要ですか?
A. Open-source, unlabeled datasets
A. オープンソースのラベルなしデータセット
B. Unlimited, publicly available internet data
B. 無制限の公開インターネットデータ
C. Strictly synthetic data
C. 厳密に合成されたデータ
D. Ethically sourced, commercially viable data
D. 倫理的に調達された商業的に実行可能なデータ
What does the abbreviation 'H × W' typically refer to in the context of image and video data resolution?
画像・動画データの解像度の文脈で「H×W」という略語は通常何を指しますか?
A. Height by Width
A. 高さ×幅
B. Horizontal by Vertical
B. 水平×垂直
C. Hue by Saturation
C. 色相×彩度
D. Height by Weight
D. 高さ×重量
What is the primary focus of Latent Diffusion Models (LDMs) in the context of image synthesis?
画像合成の文脈で潜在拡散モデル(LDM)の主な焦点は何ですか?
A. Enabling high-quality image synthesis with excessive compute demands.
A. 過剰な計算要求で高品質な画像合成を実現する
B. Focusing exclusively on low-resolution image generation.
B. 低解像度画像生成のみに集中する
C. Training a diffusion model in a compressed lower-dimensional latent space.
C. 圧縮された低次元潜在空間で拡散モデルを学習する
D. Directly generating images in pixel space without compression.
D. 圧縮なしにピクセル空間で直接画像を生成する
What kind of models currently dominate powerful image generation?
現在強力な画像生成を支配しているモデルの種類は何ですか?
A. Only autoregressive transformers.
A. 自己回帰トランスフォーマーのみ
B. Neural networks focused solely on image classification.
B. 画像分類のみに焦点を当てたニューラルネットワーク
C. A combination of generative adversarial networks, autoregressive transformers, and diffusion models.
C. GAN・自己回帰トランスフォーマー・拡散モデルの組み合わせ
D. Primarily generative adversarial networks (GANs).
D. 主に生成的敵対ネットワーク(GAN)
What is a significant challenge faced by video modeling compared to image domain advancements?
画像ドメインの進歩と比較して動画モデリングが直面する重要な課題は何ですか?
A. The low computational cost associated with training on video data.
A. 動画データ学習に伴う低い計算コスト
B. The significant computational cost and lack of extensive video datasets.
B. 大幅な計算コストと広範な動画データセットの不足
C. The ease of generating high-resolution, long videos with existing methods.
C. 既存手法での高解像度長動画生成の容易さ
D. The abundance of large-scale, general, and publicly available video datasets.
D. 大規模・汎用・公開動画データセットの豊富さ
What is one limitation of efficiently generating long videos using the approach described in Section 3.1?
セクション3.1のアプローチを使って長動画を効率的に生成する際の制限の一つは何ですか?
A. It is only suitable for generating low-resolution video sequences.
A. 低解像度の動画シーケンスのみに適している
B. It reaches its limits when synthesizing very long videos.
B. 非常に長い動画の合成時に限界に達する
C. It requires excessively high frame rates, which are difficult to achieve.
C. 達成困難な過度に高いフレームレートを必要とする
D. It eliminates the need for separate prediction models.
D. 別の予測モデルの必要性を排除する
What is the primary objective when converting an image generator into a video generator in this research?
本研究における画像生成器を動画生成器に変換する際の主な目的は何ですか?
A. To solely focus on image generation capabilities without considering temporal consistency.
A. 時間的一貫性を考慮せずに画像生成能力のみに集中する
B. To eliminate the need for pre-trained image models entirely.
B. 事前学習済み画像モデルの必要性を完全に排除する
C. To reduce the spatial resolution of the generated content.
C. 生成コンテンツの空間解像度を下げる
D. To introduce a temporal dimension to the latent space diffusion model and fine-tune on encoded image sequences.
D. 潜在空間拡散モデルに時間次元を導入し符号化画像シーケンスでFTする
Regarding the autoencoder finetuning, what is the purpose of introducing additional temporal layers for the autoencoder's decoder?
オートエンコーダのFTでデコーダに追加の時間層を導入する目的は何ですか?
A. To prevent the autoencoder from being used in latent space.
A. 潜在空間でのオートエンコーダ使用を防ぐ
B. To increase the memory requirements of the model.
B. モデルのメモリ要件を増やす
C. To counteract flickering artifacts when encoding and decoding temporally coherent image sequences.
C. 時間的に一貫した画像シーケンスのエンコード・デコード時のフリッカリングに対処する
D. To maintain frame independence and eliminate temporal awareness.
D. フレームの独立性を維持し時間的認識を排除する
What is a key contribution of this work regarding pre-trained image diffusion models?
事前学習済み画像拡散モデルに関するこの研究の重要な貢献は何ですか?
A. Eliminating the need for any diffusion models in high-resolution video synthesis.
A. 高解像度動画合成での拡散モデルの必要性を排除する
B. Focusing solely on image super-resolution without video applications.
B. 動画アプリなしに画像超解像のみに集中する
C. Leveraging pre-trained image DMs and turning them into video generators by inserting temporal layers.
C. 事前学習済み画像DMを活用し時間層を挿入して動画生成器に変換する
D. Developing entirely new diffusion models from scratch for video generation.
D. 動画生成用の全く新しい拡散モデルをゼロから開発する
What novel application is showcased for the first time using the Video LDM approach?
Video LDMアプローチを使って初めて実証された新しい応用は何ですか?
A. Creating static images with temporal layers.
A. 時間層を持つ静止画像の作成
B. Personalised text-to-video generation.
B. 個人化されたテキストから動画への生成
C. Training diffusion models with only pixel-space data.
C. ピクセル空間データのみで拡散モデルを学習する
D. Generating low-resolution videos from text prompts.
D. テキストプロンプトから低解像度動画を生成する
What is the method for converting an LDM image generator into a video generator?
LDM画像生成器を動画生成器に変換する方法は何ですか?
A. Applying the LDM solely to individual image frames sequentially.
A. LDMを個々の画像フレームに順番に適用するだけ
B. Training a new LDM entirely from scratch using video data.
B. 動画データを使って新しいLDMをゼロから学習する
C. Introducing a temporal dimension to the latent space diffusion model and fine-tuning on encoded image sequences.
C. 潜在空間拡散モデルに時間次元を導入し符号化画像シーケンスでFTする
D. Removing spatial layers and replacing them with temporal layers.
D. 空間層を除去して時間層に置き換える
How is the goal of achieving high frame rates addressed in the video synthesis process?
動画合成で高いフレームレートを達成するという目標はどのように対処されますか?
A. By relying solely on low frame rate key frames.
A. 低フレームレートのキーフレームのみに頼る
B. By introducing an additional model to interpolate between given key frames.
B. キーフレーム間を補間する追加モデルを導入する
C. By generating all frames at once in a single step.
C. 全フレームを一度に単一ステップで生成する
D. By increasing the spatial resolution instead of the temporal resolution.
D. 時間解像度の代わりに空間解像度を上げる
What does the user study on Driving Video Synthesis on RDS indicate about the 'Ours (cond.)' method compared to LVG?
RDSでの走行動画合成のユーザ研究はLVGとの比較で「Ours (cond.)」について何を示していますか?
A. Both methods achieve equal preference in realism, with no significant difference.
A. 両手法ともリアリティで同等の好みで有意差なし
B. LVG samples are generally preferred over 'Ours (cond.)' in terms of realism.
B. LVGサンプルがOurs (cond.)より一般的に好まれる
C. The 'Ours (cond.)' method is preferred in terms of realism by 8% of participants over LVG.
C. Ours (cond.)はLVGより8%の参加者にリアリティで好まれる
D. Samples from 'Ours (cond.)' are preferred over unconditional samples and LVG samples in terms of realism.
D. Ours (cond.)のサンプルが無条件サンプルとLVGより好まれる
What was the observed impact on FVD and FID when video frames were upsampled independently, rather than with temporal alignment of the upsampler?
時間的整合性なしにフレームを独立してアップサンプリングした場合のFVDとFIDへの影響は何ですか?
A. FVD degraded significantly while FID remained essentially unaffected, showing loss of temporal consistency.
A. FVDは大幅に劣化したがFIDは基本的に影響なく時間的一貫性の損失を示した
B. Neither FVD nor FID showed any change, suggesting no impact from temporal alignment.
B. FVDもFIDも変化なく時間的整合性の影響なし
C. FVD improved significantly, while FID degraded, indicating better temporal consistency but worse image quality.
C. FVDは大幅に改善したがFIDは劣化した
D. Both FVD and FID improved significantly, indicating better performance.
D. FVDとFIDの両方が大幅に改善した
What kind of scenarios can a high-resolution video generator trained on in-the-wild driving scenes simulate?
野生の走行シーンで学習された高解像度動画生成器はどのようなシナリオをシミュレートできますか?
A. Several different plausible future predictions given an initial frame.
A. 初期フレームが与えられた複数の異なる可能な未来予測
B. Only low-resolution, short driving clips.
B. 低解像度の短い走行クリップのみ
C. Only predetermined, fixed driving scenarios.
C. あらかじめ決められた固定の走行シナリオのみ
D. Static images of driving environments without motion.
D. 動きのない走行環境の静止画像
How does the approach enable a user to create customized driving scenario simulations?
このアプローチはユーザーがカスタマイズされた走行シナリオシミュレーションを作成するためにどのように機能しますか?
A. By providing pre-generated video templates for common scenarios.
A. 一般的なシナリオ用の事前生成動画テンプレートを提供する
B. By manually creating a scene composition through specifying bounding boxes for cars and using this as an initialization.
B. 車のバウンディングボックスを指定してシーン構成を手動で作成し初期化として使用する
C. By only using existing images from the dataset without any user input.
C. ユーザー入力なしにデータセットの既存画像のみを使用する
D. By training a new Video LDM for each desired scenario from scratch.
D. 望ましいシナリオごとに新しいVideo LDMをゼロから学習する
What are two real-world applications focused on in this research?
本研究で焦点を当てている2つの実世界の応用は何ですか?
A. Simulation of in-the-wild driving data and creative content creation with text-to-video modeling.
A. 野生の走行データのシミュレーションとテキストから動画へのモデリングによるコンテンツ制作
B. Weather forecasting and astronomical simulations.
B. 気象予報と天文シミュレーション
C. Financial market prediction and natural language processing.
C. 金融市場予測と自然言語処理
D. Medical imaging analysis and robotics control.
D. 医療画像解析とロボティクス制御
How does the Video LDM contribute to personalized text-to-video generation?
Video LDMはどのようにして個人化されたテキストから動画への生成に貢献しますか?
A. By limiting generation to generic, non-personalized content.
A. 汎用的な非個人化コンテンツのみに生成を制限する
B. By altering the spatial layers for each personalized request.
B. 個人化リクエストごとに空間層を変更する
C. By transferring temporal layers trained on one Image LDM backbone to other fine-tuned model checkpoints.
C. ある画像LDMバックボーンで学習した時間層を他のFTモデルチェックポイントに転送する
D. By requiring extensive new training data for each personalized item.
D. 個人化された各アイテムに大量の新しい学習データを必要とする
What is a stated limitation of the current synthesized videos generated by the Video LDM?
Video LDMが生成した現在の合成動画の述べられた限界は何ですか?
A. They can only be generated in low resolution.
A. 低解像度でのみ生成できる
B. They lack any ethical or safety implications.
B. 倫理的・安全性への影響が全くない
C. They are not indistinguishable from real content yet.
C. まだリアルなコンテンツと区別がつかない状態ではない
D. They are indistinguishable from real content.
D. リアルなコンテンツと区別がつかない
What is an important direction for future work regarding large-scale generative models?
大規模生成モデルに関する将来の研究の重要な方向性は何ですか?
A. Reducing model size to limit performance.
A. パフォーマンスを制限するためにモデルサイズを縮小する
B. Avoiding any ethical considerations in model development.
B. モデル開発での倫理的考慮を避ける
C. Focusing exclusively on commercially non-viable data.
C. 商業的に実行不可能なデータのみに集中する
D. Training with ethically sourced, commercially viable data.
D. 倫理的に調達された商業的に実行可能なデータで学習する
What is the key design choice of the presented Video Latent Diffusion Models (Video LDMs) for high-resolution video generation?
高解像度動画生成のためのVideo LDMsの主要な設計選択は何ですか?
A. Building on pre-trained image diffusion models and temporally video fine-tuning them with alignment layers.
A. 事前学習済み画像拡散モデルを基に整合層で時間的動画FTを行う
B. Building new video models entirely from scratch without leveraging existing image models.
B. 既存の画像モデルを活用せずに全く新しい動画モデルをゼロから構築する
C. Relying solely on pixel-space diffusion models for all video generation tasks.
C. 全動画生成タスクにピクセル空間拡散モデルのみを使用する
D. Prioritizing computational intensity to ensure high resolution.
D. 高解像度を確保するために計算強度を優先する
What is a main conclusion regarding the learned temporal layers in the Video LDM?
Video LDMで学習された時間層に関する主な結論は何ですか?
A. They transfer to different model checkpoints, enabling personalized text-to-video generation.
A. 異なるモデルチェックポイントに転送可能で個人化テキストから動画への生成を可能にする
B. They are highly specific to the initial model and do not generalize.
B. 初期モデルに非常に特有で汎化しない
C. They increase computational cost without significant benefits.
C. 大幅な利益なしに計算コストを増加させる
D. They only work with low-resolution video generation tasks.
D. 低解像度の動画生成タスクにのみ機能する
Why has video modeling lagged behind image domain advancements?
動画モデリングはなぜ画像ドメインの進歩に遅れをとっているのですか?
A. Mainly due to the significant computational cost associated with training on video data, and the lack of large-scale, general, and publicly available video datasets.
A. 主に動画データ学習の大幅な計算コストと大規模・汎用・公開動画データセットの不足のため
B. Due to insufficient computational cost associated with video data.
B. 動画データに伴う不十分な計算コストのため
C. Because there are no real-world problems that require video generation.
C. 動画生成を必要とする現実世界の問題がないため
D. Because video data is too easy to acquire and process.
D. 動画データの取得と処理が容易すぎるため
How does the approach leverage off-the-shelf pre-trained image LDMs efficiently?
このアプローチは市販の事前学習済み画像LDMをどのように効率的に活用しますか?
A. By discarding the spatial layers of the pre-trained image LDM.
A. 事前学習済み画像LDMの空間層を廃棄することによって
B. By retraining the entire LDM from scratch for video tasks.
B. 動画タスクのためにLDM全体をゼロから再学習することによって
C. By only needing to train a temporal alignment model in that case.
C. その場合に時間的整合モデルのみを学習することが必要なことによって
D. By using an entirely different architecture for video generation.
D. 動画生成に全く異なるアーキテクチャを使用することによって
What physical constraint often dictates that the video upsampler only needs to operate locally for synthesis at very high resolutions?
非常に高い解像度での合成において動画アップサンプラーがローカルにのみ動作する必要がある物理的制約は何ですか?
A. The absence of long-term temporal correlations in upscaling tasks.
A. アップスケーリングタスクにおける長期的時間相関の欠如
B. The inherent difficulty in processing global information at high resolutions.
B. 高解像度でのグローバル情報処理の固有の難しさ
C. The limited capabilities of pixel-space diffusion models.
C. ピクセル空間拡散モデルの限られた能力
D. The need to maintain low training and computational requirements.
D. 低い学習・計算要件を維持する必要性
According to the human evaluation on driving video synthesis, how do 'Ours (cond.)' samples compare to 'Ours (uncond.)' samples in terms of realism?
走行動画合成の人間評価によると「Ours (cond.)」は「Ours (uncond.)」とリアリティの観点でどのように比較されますか?
A. The human evaluation showed no preference between the two methods.
A. 人間評価では2手法間の優位性の差はなかった
B. Both methods are equally preferred, indicating no significant difference.
B. 両手法は同等に好まれ有意差なし
C. 'Ours (cond.)' samples are preferred over 'Ours (uncond.)' samples by about 7% of participants.
C. Ours (cond.)のサンプルはOurs (uncond.)より約7%の参加者に好まれる
D. 'Ours (uncond.)' samples are largely preferred over 'Ours (cond.)' samples.
D. Ours (uncond.)のサンプルはOurs (cond.)より大幅に好まれる
What is a potential negative implication of enhanced versions of the Video LDM reaching even higher quality?
Video LDMの強化版がより高品質に達することの潜在的な否定的含意は何ですか?
A. It might increase the cost of video content creation for digital artists.
A. デジタルアーティストの動画コンテンツ制作コストが増加する可能性
B. It could generate videos that appear deceptively real, leading to ethical and safety implications due to malicious use.
B. 欺瞞的にリアルに見える動画を生成し悪意ある使用による倫理的・安全性への影響をもたらす可能性
C. It could limit artistic expression to only a few individuals.
C. 芸術的表現を少数の個人のみに制限する可能性
D. It may reduce the need for autonomous driving simulators.
D. 自律走行シミュレータの必要性を減らす可能性
(参照: Y. Chen et al., "Paper2Web: Let's Make Your Paper Alive!," arXiv:2510.15842, 2025. https://arxiv.org/abs/2510.15842)
定性評価
同じ論文でも,Project Pageの作り方で伝わりやすさは変わるのか?
Project Page のわかりやすさに寄与するデザイン要因を明らかにするため,CVPR 2023–2025 の高被引用論文 75 本の Project Page を 6 名で精読し,"わかりやすさ" を左右する特徴を「第一印象」「操作性」「図と文の比率」「手法の伝達」の 4 観点に整理した. ここでは代表例として,CVPR 2025 に Oral として採択された論文 "Continuous 3D Perception Model with Persistent State"(CUT3R,Wang et al.)の Project Page を取り上げ,実際に公開されている Project Page(A. Original),Paper2Web による自動生成例(B. Paper2Web),悪い Project Page の特徴をもとに Claude Code で生成した例(C. Bad example)を 4 観点で比較する.
3 例の比較から,同じ論文であっても Project Page の設計次第で伝わりやすさが大きく変わることが確認できた.冒頭の動画と触れるデモを備えた A. Original に対し,B. Paper2Web は体裁こそ整っているものの図が少なく解説が冗長で,C. Bad example は文字のみで内容の判読が困難であった.
以上より,冒頭で内容が伝わる第一印象,動かして確かめられる操作性,図で見せる図と文章のバランスが,わかりやすい Project Page の条件として重要である.Project Page は論文では示せない動画や補足情報を発信できる媒体であり,こうした特性を活かした設計が研究の理解促進に寄与すると考えられる.
(参照: Q. Wang et al., "Continuous 3D Perception Model with Persistent State," CVPR 2025. https://cut3r.github.io/)
謝辞
本研究はMIRU2026 若手プログラムにおける活動の一環として行われました.また,本プログラムを通してメンターとしてご指導いただいた上田栞さんに深く感謝いたします.
BibTeX
@misc{MIRUwakate2026projectpage,
title = {わかりやすいプロジェクトページを求めて},
author = {Arimura, Koshun and
Oumi, Yusuke and
Konno, Tsubasa and
Sugaya, Tomoki and
Takezaki, Shumpei and
Yamamoto, Riku},
howpublished = {MIRU2026 若手プログラム},
year = {2026}
}
わかりやすいプロジェクトページを求めて