Awesome VLM Architectures
VLM Architecturesを扱う資料や関連プロジェクトをまとめたAwesomeリストです。
目次
引用
このリポジトリが役立つ場合は、以下の情報で引用できます。個別モデルの主張には原論文も引用してください。このカタログは文献案内であり、原資料の代替ではありません。
Gökay Aydoğanがfal.aiで作成・保守しています(ORCID、gokay@fal.ai)。
構成図は図版クレジットで個別に帰属を示し、図版通知記載の権利に従います。
📚 BibTeX
@misc{aydogan2024awesomevlmarchitectures,
author = {Gökay Aydoğan},
title = {Awesome VLM Architectures},
year = {2024},
howpublished = {\url{https://github.com/gokayfem/awesome-vlm-architectures}},
note = {GitHub repository, fal.ai},
url = {https://github.com/gokayfem/awesome-vlm-architectures}
}モデル
すべてのアーキテクチャパネルはリリース日の新しい順です。同日公開のモデルは編集上のカタログ順を維持します。
🧭 Chronological Model Index (155 architectures, newest first)
2026: MODUS | Argus-Unified | Kimi K3 | Mage-VL | Inkling | Hy-Embodied-VLM | MonkeyOCRv2 | MiniMax M3 | InternVideo3 | Keye-VL 2.0 | Zamba2-VL | Cosmos 3 | Lance | ZAYA1-VL | Falcon Perception | GLM-5V-Turbo | PLaMo 2.1-VL | EXAONE 4.5 | BidirLM and BidirLM-Omni | Gemma 4 | Penguin-VL | Phi-4-Reasoning-Vision | V-SONAR and V-LCM | Qwen3.5 | Youtu-VL | Kimi K2.5 and K2.6 | Step3-VL-10B
2025: ERNIE 5.0 | DeepSeek-OCR | PaddleOCR-VL | Qwen3-VL | Step3 | GLM-4.1V-Thinking | ERNIE 4.5-VL | MiMo-VL | BAGEL | Seed1.5-VL | InternVL3 and InternVL3.5 | Kimi-VL | Llama 4 Scout and Maverick | Qwen2.5-Omni | Gemma 3 | Aya Vision | Phi-4-multimodal | SigLIP 2 | EVEv2 | Qwen2.5-VL | VideoLLaMA 3 | UI-TARS | MiniMax-01 | MiniCPM-o-2.6 | Eagle 2 | Sa2VA
2024: VideoChat-Flash | OmniVLM | Apollo | DeepSeek-VL2 | Maya | InternVL 2.5 | PaliGemma 2 | ShowUI | SmolVLM | AIMv2 | LLaVA-CoT | LLM2CLIP | Tarsier2 | Janus and Janus-Pro | ARIA | Emu3 | Molmo and PixMo | Llama 3.2-Vision | NVLM | Pixtral 12B | VILA-U | Qwen2-VL | EAGLE | Show-o | Idefics3-8B | Transfusion | mPLUG-Owl3 | VITA | LLaVA-OneVision | VILA² | INF-LLaVA | SlowFast-LLaVA | EVLM | InternLM-XComposer-2.5 | OMG-LLaVA | Cambrian-1 | EVE | Ovis | Parrot | ConvLLaVA | Phi-3-Vision and Phi-3.5-Vision | CogVLM2 | Chameleon | PaliGemma | xGen-MM (BLIP-3) | MANTIS | Moondream-next | Idefics2 | InternLM-XComposer2-4KHD | MM1 | DeepSeek-VL | AnyGPT | SPHINX-X | LLaVA 1.6 | MiniCPM-V | MouSi | InternLM-XComposer2 | MoE-LLaVA | moondream1 and moondream2 | FireLLaVA | COSMO
2023: TinyGPT-V | MobileVLM | Alpha-CLIP | Nous-Hermes-2-Vision - Mistral 7B | SPHINX | Florence-2 | u-LLaVA | LLaVA-Plus | OtterHD | CoVLM | GLaMM | Fuyu-8B | PaLI-3 Vision Language Models | MiniGPT-v2 | BakLLaVA | Ferret | LLaVA 1.5 | CogVLM | MetaCLIP | Qwen-VL | IDEFICS | BLIVA | KOSMOS-2 | LaVIN | InstructBLIP | ImageBind | LLaVA | MiniGPT-4 | SigLIP | OpenFlamingo | PaLM-E | KOSMOS-1 | BLIP-2
2022: MULTIINSTRUCT | PaLI | Flamingo | BLIP
2020: ViT
リリース年表
日付は確認できる最初の公式モデルリリースを使い、ない場合は論文のarXiv v1投稿または最初の技術報告を使います。同系列のポイントリリースは最初のアーキテクチャ公開へ統合し、同日項目はカタログ順を維持します。
🗓️ Release Timeline (155 architectures, newest first)
| 日付 | アーキテクチャ | 特徴的な貢献 |
|---|---|---|
| 2026-07-28 | MODUS | MODUSの特徴的なアーキテクチャ上の貢献 |
| 2026-07-28 | Argus-Unified | Argus-Unifiedの特徴的なアーキテクチャ上の貢献 |
| 2026-07-27 | Kimi K3 | Kimi K3の特徴的なアーキテクチャ上の貢献 |
| 2026-07-27 | Mage-VL | Mage-VLの特徴的なアーキテクチャ上の貢献 |
| 2026-07-15 | Inkling | Inklingの特徴的なアーキテクチャ上の貢献 |
| 2026-07-15 | Hy-Embodied-VLM | Hy-Embodied-VLMの特徴的なアーキテクチャ上の貢献 |
| 2026-07-11 | MonkeyOCRv2 | MonkeyOCRv2の特徴的なアーキテクチャ上の貢献 |
| 2026-06-11 | MiniMax M3 | MiniMax M3の特徴的なアーキテクチャ上の貢献 |
| 2026-06-10 | InternVideo3 | InternVideo3の特徴的なアーキテクチャ上の貢献 |
| 2026-06-09 | Keye-VL 2.0 | Keye-VL 2.0の特徴的なアーキテクチャ上の貢献 |
| 2026-06-02 | Zamba2-VL | Zamba2-VLの特徴的なアーキテクチャ上の貢献 |
| 2026-05-31 | Cosmos 3 | Cosmos 3の特徴的なアーキテクチャ上の貢献 |
| 2026-05-18 | Lance | Lanceの特徴的なアーキテクチャ上の貢献 |
| 2026-05-08 | ZAYA1-VL | ZAYA1-VLの特徴的なアーキテクチャ上の貢献 |
| 2026-05-03 | Falcon Perception | Falcon Perceptionの特徴的なアーキテクチャ上の貢献 |
| 2026-04-29 | GLM-5V-Turbo | GLM-5V-Turboの特徴的なアーキテクチャ上の貢献 |
| 2026-04-21 | PLaMo 2.1-VL | PLaMo 2.1-VLの特徴的なアーキテクチャ上の貢献 |
| 2026-04-09 | EXAONE 4.5 | EXAONE 4.5の特徴的なアーキテクチャ上の貢献 |
| 2026-04-02 | BidirLM and BidirLM-Omni | BidirLM and BidirLM-Omniの特徴的なアーキテクチャ上の貢献 |
| 2026-03-31 | Gemma 4 | Gemma 4の特徴的なアーキテクチャ上の貢献 |
| 2026-03-06 | Penguin-VL | Penguin-VLの特徴的なアーキテクチャ上の貢献 |
| 2026-03-04 | Phi-4-Reasoning-Vision | Phi-4-Reasoning-Visionの特徴的なアーキテクチャ上の貢献 |
| 2026-03-01 | V-SONAR and V-LCM | V-SONAR and V-LCMの特徴的なアーキテクチャ上の貢献 |
| 2026-02-16 | Qwen3.5 | Qwen3.5の特徴的なアーキテクチャ上の貢献 |
| 2026-01-27 | Youtu-VL | Youtu-VLの特徴的なアーキテクチャ上の貢献 |
| 2026-01-27 | Kimi K2.5 and K2.6 | Kimi K2.5 and K2.6の特徴的なアーキテクチャ上の貢献 |
| 2026-01-14 | Step3-VL-10B | Step3-VL-10Bの特徴的なアーキテクチャ上の貢献 |
| 2025-11-13 | ERNIE 5.0 | ERNIE 5.0の特徴的なアーキテクチャ上の貢献 |
| 2025-10-20 | DeepSeek-OCR | DeepSeek-OCRの特徴的なアーキテクチャ上の貢献 |
| 2025-10-16 | PaddleOCR-VL | PaddleOCR-VLの特徴的なアーキテクチャ上の貢献 |
| 2025-09-22 | Qwen3-VL | Qwen3-VLの特徴的なアーキテクチャ上の貢献 |
| 2025-07-25 | Step3 | Step3の特徴的なアーキテクチャ上の貢献 |
| 2025-07-01 | GLM-4.1V-Thinking | GLM-4.1V-Thinkingの特徴的なアーキテクチャ上の貢献 |
| 2025-06-30 | ERNIE 4.5-VL | ERNIE 4.5-VLの特徴的なアーキテクチャ上の貢献 |
| 2025-06-04 | MiMo-VL | MiMo-VLの特徴的なアーキテクチャ上の貢献 |
| 2025-05-20 | BAGEL | BAGELの特徴的なアーキテクチャ上の貢献 |
| 2025-05-11 | Seed1.5-VL | Seed1.5-VLの特徴的なアーキテクチャ上の貢献 |
| 2025-04-11 | InternVL3 and InternVL3.5 | InternVL3 and InternVL3.5の特徴的なアーキテクチャ上の貢献 |
| 2025-04-10 | Kimi-VL | Kimi-VLの特徴的なアーキテクチャ上の貢献 |
| 2025-04-05 | Llama 4 Scout and Maverick | Llama 4 Scout and Maverickの特徴的なアーキテクチャ上の貢献 |
| 2025-03-26 | Qwen2.5-Omni | Qwen2.5-Omniの特徴的なアーキテクチャ上の貢献 |
| 2025-03-12 | Gemma 3 | Gemma 3の特徴的なアーキテクチャ上の貢献 |
| 2025-03-04 | Aya Vision | Aya Visionの特徴的なアーキテクチャ上の貢献 |
| 2025-03-03 | Phi-4-multimodal | Phi-4-multimodalの特徴的なアーキテクチャ上の貢献 |
| 2025-02-20 | SigLIP 2 | SigLIP 2の特徴的なアーキテクチャ上の貢献 |
| 2025-02-08 | EVEv2 | EVEv2の特徴的なアーキテクチャ上の貢献 |
| 2025-01-26 | Qwen2.5-VL | Qwen2.5-VLの特徴的なアーキテクチャ上の貢献 |
| 2025-01-21 | VideoLLaMA 3 | VideoLLaMA 3の特徴的なアーキテクチャ上の貢献 |
| 2025-01-20 | UI-TARS | UI-TARSの特徴的なアーキテクチャ上の貢献 |
| 2025-01-14 | MiniMax-01 | MiniMax-01の特徴的なアーキテクチャ上の貢献 |
| 2025-01-12 | MiniCPM-o-2.6 | MiniCPM-o-2.6の特徴的なアーキテクチャ上の貢献 |
| 2025-01-10 | Eagle 2 | Eagle 2の特徴的なアーキテクチャ上の貢献 |
| 2025-01-07 | Sa2VA | Sa2VAの特徴的なアーキテクチャ上の貢献 |
| 2024-12-31 | VideoChat-Flash | VideoChat-Flashの特徴的なアーキテクチャ上の貢献 |
| 2024-12-16 | OmniVLM | OmniVLMの特徴的なアーキテクチャ上の貢献 |
| 2024-12-13 | Apollo | Apolloの特徴的なアーキテクチャ上の貢献 |
| 2024-12-13 | DeepSeek-VL2 | DeepSeek-VL2の特徴的なアーキテクチャ上の貢献 |
| 2024-12-10 | Maya | Mayaの特徴的なアーキテクチャ上の貢献 |
| 2024-12-05 | InternVL 2.5 | InternVL 2.5の特徴的なアーキテクチャ上の貢献 |
| 2024-12-04 | PaliGemma 2 | PaliGemma 2の特徴的なアーキテクチャ上の貢献 |
| 2024-11-26 | ShowUI | ShowUIの特徴的なアーキテクチャ上の貢献 |
| 2024-11-26 | SmolVLM | SmolVLMの特徴的なアーキテクチャ上の貢献 |
| 2024-11-21 | AIMv2 | AIMv2の特徴的なアーキテクチャ上の貢献 |
| 2024-11-15 | LLaVA-CoT | LLaVA-CoTの特徴的なアーキテクチャ上の貢献 |
| 2024-11-06 | LLM2CLIP | LLM2CLIPの特徴的なアーキテクチャ上の貢献 |
| 2024-11-05 | Tarsier2 | Tarsier2の特徴的なアーキテクチャ上の貢献 |
| 2024-10-17 | Janus and Janus-Pro | Janus and Janus-Proの特徴的なアーキテクチャ上の貢献 |
| 2024-10-08 | ARIA | ARIAの特徴的なアーキテクチャ上の貢献 |
| 2024-09-27 | Emu3 | Emu3の特徴的なアーキテクチャ上の貢献 |
| 2024-09-25 | Molmo and PixMo | Molmo and PixMoの特徴的なアーキテクチャ上の貢献 |
| 2024-09-25 | Llama 3.2-Vision | Llama 3.2-Visionの特徴的なアーキテクチャ上の貢献 |
| 2024-09-17 | NVLM | NVLMの特徴的なアーキテクチャ上の貢献 |
| 2024-09-11 | Pixtral 12B | Pixtral 12Bの特徴的なアーキテクチャ上の貢献 |
| 2024-09-06 | VILA-U | VILA-Uの特徴的なアーキテクチャ上の貢献 |
| 2024-08-29 | Qwen2-VL | Qwen2-VLの特徴的なアーキテクチャ上の貢献 |
| 2024-08-28 | EAGLE | EAGLEの特徴的なアーキテクチャ上の貢献 |
| 2024-08-22 | Show-o | Show-oの特徴的なアーキテクチャ上の貢献 |
| 2024-08-22 | Idefics3-8B | Idefics3-8Bの特徴的なアーキテクチャ上の貢献 |
| 2024-08-20 | Transfusion | Transfusionの特徴的なアーキテクチャ上の貢献 |
| 2024-08-09 | mPLUG-Owl3 | mPLUG-Owl3の特徴的なアーキテクチャ上の貢献 |
| 2024-08-09 | VITA | VITAの特徴的なアーキテクチャ上の貢献 |
| 2024-08-05 | LLaVA-OneVision | LLaVA-OneVisionの特徴的なアーキテクチャ上の貢献 |
| 2024-07-24 | VILA² | VILA²の特徴的なアーキテクチャ上の貢献 |
| 2024-07-23 | INF-LLaVA | INF-LLaVAの特徴的なアーキテクチャ上の貢献 |
| 2024-07-22 | SlowFast-LLaVA | SlowFast-LLaVAの特徴的なアーキテクチャ上の貢献 |
| 2024-07-19 | EVLM | EVLMの特徴的なアーキテクチャ上の貢献 |
| 2024-07-03 | InternLM-XComposer-2.5 | InternLM-XComposer-2.5の特徴的なアーキテクチャ上の貢献 |
| 2024-06-27 | OMG-LLaVA | OMG-LLaVAの特徴的なアーキテクチャ上の貢献 |
| 2024-06-24 | Cambrian-1 | Cambrian-1の特徴的なアーキテクチャ上の貢献 |
| 2024-06-17 | EVE | EVEの特徴的なアーキテクチャ上の貢献 |
| 2024-06-14 | Ovis | Ovisの特徴的なアーキテクチャ上の貢献 |
| 2024-06-04 | Parrot | Parrotの特徴的なアーキテクチャ上の貢献 |
| 2024-05-24 | ConvLLaVA | ConvLLaVAの特徴的なアーキテクチャ上の貢献 |
| 2024-05-21 | Phi-3-Vision and Phi-3.5-Vision | Phi-3-Vision and Phi-3.5-Visionの特徴的なアーキテクチャ上の貢献 |
| 2024-05-20 | CogVLM2 | CogVLM2の特徴的なアーキテクチャ上の貢献 |
| 2024-05-16 | Chameleon | Chameleonの特徴的なアーキテクチャ上の貢献 |
| 2024-05-14 | PaliGemma | PaliGemmaの特徴的なアーキテクチャ上の貢献 |
| 2024-05-06 | xGen-MM (BLIP-3) | xGen-MM (BLIP-3)の特徴的なアーキテクチャ上の貢献 |
| 2024-05-02 | MANTIS | MANTISの特徴的なアーキテクチャ上の貢献 |
| 2024-04-19 | Moondream-next | Moondream-nextの特徴的なアーキテクチャ上の貢献 |
| 2024-04-15 | Idefics2 | Idefics2の特徴的なアーキテクチャ上の貢献 |
| 2024-04-09 | InternLM-XComposer2-4KHD | InternLM-XComposer2-4KHDの特徴的なアーキテクチャ上の貢献 |
| 2024-03-14 | MM1 | MM1の特徴的なアーキテクチャ上の貢献 |
| 2024-03-08 | DeepSeek-VL | DeepSeek-VLの特徴的なアーキテクチャ上の貢献 |
| 2024-02-19 | AnyGPT | AnyGPTの特徴的なアーキテクチャ上の貢献 |
| 2024-02-08 | SPHINX-X | SPHINX-Xの特徴的なアーキテクチャ上の貢献 |
| 2024-01-30 | LLaVA 1.6 | LLaVA 1.6の特徴的なアーキテクチャ上の貢献 |
| 2024-01-30 | MiniCPM-V | MiniCPM-Vの特徴的なアーキテクチャ上の貢献 |
| 2024-01-30 | MouSi | MouSiの特徴的なアーキテクチャ上の貢献 |
| 2024-01-29 | InternLM-XComposer2 | InternLM-XComposer2の特徴的なアーキテクチャ上の貢献 |
| 2024-01-29 | MoE-LLaVA | MoE-LLaVAの特徴的なアーキテクチャ上の貢献 |
| 2024-01-20 | moondream1 and moondream2 | moondream1 and moondream2の特徴的なアーキテクチャ上の貢献 |
| 2024-01-05 | FireLLaVA | FireLLaVAの特徴的なアーキテクチャ上の貢献 |
| 2024-01-01 | COSMO | COSMOの特徴的なアーキテクチャ上の貢献 |
| 2023-12-28 | TinyGPT-V | TinyGPT-Vの特徴的なアーキテクチャ上の貢献 |
| 2023-12-28 | MobileVLM | MobileVLMの特徴的なアーキテクチャ上の貢献 |
| 2023-12-06 | Alpha-CLIP | Alpha-CLIPの特徴的なアーキテクチャ上の貢献 |
| 2023-11-28 | Nous-Hermes-2-Vision - Mistral 7B | Nous-Hermes-2-Vision - Mistral 7Bの特徴的なアーキテクチャ上の貢献 |
| 2023-11-13 | SPHINX | SPHINXの特徴的なアーキテクチャ上の貢献 |
| 2023-11-10 | Florence-2 | Florence-2の特徴的なアーキテクチャ上の貢献 |
| 2023-11-09 | u-LLaVA | u-LLaVAの特徴的なアーキテクチャ上の貢献 |
| 2023-11-09 | LLaVA-Plus | LLaVA-Plusの特徴的なアーキテクチャ上の貢献 |
| 2023-11-07 | OtterHD | OtterHDの特徴的なアーキテクチャ上の貢献 |
| 2023-11-06 | CoVLM | CoVLMの特徴的なアーキテクチャ上の貢献 |
| 2023-11-06 | GLaMM | GLaMMの特徴的なアーキテクチャ上の貢献 |
| 2023-10-17 | Fuyu-8B | Fuyu-8Bの特徴的なアーキテクチャ上の貢献 |
| 2023-10-13 | PaLI-3 Vision Language Models | PaLI-3 Vision Language Modelsの特徴的なアーキテクチャ上の貢献 |
| 2023-10-13 | MiniGPT-v2 | MiniGPT-v2の特徴的なアーキテクチャ上の貢献 |
| 2023-10-12 | BakLLaVA | BakLLaVAの特徴的なアーキテクチャ上の貢献 |
| 2023-10-11 | Ferret | Ferretの特徴的なアーキテクチャ上の貢献 |
| 2023-10-05 | LLaVA 1.5 | LLaVA 1.5の特徴的なアーキテクチャ上の貢献 |
| 2023-10-05 | CogVLM | CogVLMの特徴的なアーキテクチャ上の貢献 |
| 2023-09-28 | MetaCLIP | MetaCLIPの特徴的なアーキテクチャ上の貢献 |
| 2023-08-24 | Qwen-VL | Qwen-VLの特徴的なアーキテクチャ上の貢献 |
| 2023-08-22 | IDEFICS | IDEFICSの特徴的なアーキテクチャ上の貢献 |
| 2023-08-19 | BLIVA | BLIVAの特徴的なアーキテクチャ上の貢献 |
| 2023-06-26 | KOSMOS-2 | KOSMOS-2の特徴的なアーキテクチャ上の貢献 |
| 2023-05-24 | LaVIN | LaVINの特徴的なアーキテクチャ上の貢献 |
| 2023-05-11 | InstructBLIP | InstructBLIPの特徴的なアーキテクチャ上の貢献 |
| 2023-05-09 | ImageBind | ImageBindの特徴的なアーキテクチャ上の貢献 |
| 2023-04-17 | LLaVA | LLaVAの特徴的なアーキテクチャ上の貢献 |
| 2023-04-16 | MiniGPT-4 | MiniGPT-4の特徴的なアーキテクチャ上の貢献 |
| 2023-03-27 | SigLIP | SigLIPの特徴的なアーキテクチャ上の貢献 |
| 2023-03-14 | OpenFlamingo | OpenFlamingoの特徴的なアーキテクチャ上の貢献 |
| 2023-03-06 | PaLM-E | PaLM-Eの特徴的なアーキテクチャ上の貢献 |
| 2023-02-27 | KOSMOS-1 | KOSMOS-1の特徴的なアーキテクチャ上の貢献 |
| 2023-01-30 | BLIP-2 | BLIP-2の特徴的なアーキテクチャ上の貢献 |
| 2022-12-21 | MULTIINSTRUCT | MULTIINSTRUCTの特徴的なアーキテクチャ上の貢献 |
| 2022-09-14 | PaLI | PaLIの特徴的なアーキテクチャ上の貢献 |
| 2022-04-28 | Flamingo | Flamingoの特徴的なアーキテクチャ上の貢献 |
| 2022-01-28 | BLIP | BLIPの特徴的なアーキテクチャ上の貢献 |
| 2021-12-07 | GLIP | GLIPの特徴的なアーキテクチャ上の貢献 |
| 2021-06-25 | FROZEN | FROZENの特徴的なアーキテクチャ上の貢献 |
| 2021-01-05 | CLIP | CLIPの特徴的なアーキテクチャ上の貢献 |
| 2020-10-22 | ViT | ViTの特徴的なアーキテクチャ上の貢献 |
アーキテクチャ
MODUS: Decoder-Only Any-to-Any Multimodal Modeling
MODUSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Mingqiao Ye et al., EPFL
Released: 2026-07-28
Figure 2. Decoder-only any-to-any modeling across tokenized 1D and 2D modalities. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
MODUSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
MODUSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Argus-Unified: Economical Understanding and Generation
Argus-Unifiedの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Weiming Zhuang et al.
Released: 2026-07-28
Figure 3. Two-stage hybrid-token training for unified image understanding and generation. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
Argus-Unifiedの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Argus-Unifiedの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 15.6、2,000。
Kimi K3: Kimi Delta Attention at Trillion-Parameter Scale
Kimi K3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 896。
Kimi Team, Moonshot AI
Released: 2026-07-27
Figure 2. Kimi Delta Attention, Stable LatentMoE, Attention Residuals, and MoonViT-V2. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
Kimi K3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.8T、104B、93、24、16、896。
Kimi K3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 401M、1,048,576-。
Mage-VL: Codec-Native Streaming Multimodality
Mage-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Senqiao Yang et al., Microsoft Research
Released: 2026-07-27
Figure 3. Codec-native streaming perception with an event gate and causal language decoder. Source paper, PDF p. 8. Figure notice.
ℹ️ 詳細情報
Mage-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 16×16-、75、560M、100M。
Mage-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3.5。
Inkling: Relative-Position Multimodal Mixture of Experts
Inklingの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Thinking Machines Lab
Released: 2026-07-15
Architecture figure: The official Inkling model card contains no architecture figure.
ℹ️ 詳細情報
Inklingの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 975B、41B、256。
Inklingの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 45T。
Hy-Embodied-VLM: Sparse-MoE Reasoning for Physical Agents
Hy-Embodied-VLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.0。
Tencent Robotics X, Hy Vision Team and Futian Laboratory
Released: 2026-07-15
Figure 4. Self-evolving supervised fine-tuning, rejection sampling, and specialized reinforcement learning. Source paper, PDF p. 11. Figure notice.
ℹ️ 詳細情報
Hy-Embodied-VLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.0、30B、3B、128、32K。
Hy-Embodied-VLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
MonkeyOCRv2: Document-Native Visual-Text Pretraining
MonkeyOCRv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 17。
Yuliang Liu et al.
Released: 2026-07-11
Figure 1. Document-native pretraining through text generation and pixel reconstruction. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
MonkeyOCRv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
MonkeyOCRv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 113M、17、0.7B、11。
MiniMax M3: Native Multimodality with Sparse Long-Context Attention
MiniMax M3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 109B。
MiniMax
Released: 2026-06-11
Figure 1. MiniMax Sparse Attention index and exact-attention branches. Source paper, PDF p. 1. Figure notice.
ℹ️ 詳細情報
MiniMax M3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
MiniMax M3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 109B。
InternVideo3: Multimodal Contextual Reasoning for Video Agents
InternVideo3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Ziang Yan et al.
Released: 2026-06-10
Figure 2. InternVideo3 with multimodal multi-head latent attention across long contexts. Source paper, PDF p. 7. Figure notice.
ℹ️ 詳細情報
InternVideo3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
InternVideo3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Keye-VL 2.0: Sparse Attention for Long-Video Agents
Keye-VL 2.0の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.0、256K。
Kwai Keye Team
Released: 2026-06-09
Figure 2. Four-stage curriculum extending Keye-VL from alignment to 256K context. Source paper, PDF p. 8. Figure notice.
ℹ️ 詳細情報
Keye-VL 2.0の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.0-30B、3B、30B、256K。
Keye-VL 2.0の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Zamba2-VL: Hybrid State-Space Vision-Language Modeling
Zamba2-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。
Zyphra
Released: 2026-06-02
Figure 1. Zamba2 hybrid state-space language backbone connected to a vision encoder. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
Zamba2-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.2B、2.7B、7B、2。
Zamba2-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Cosmos 3: Omnimodal World Modeling with Mixture of Transformers
Cosmos 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3。
NVIDIA
Released: 2026-05-31
Figure 5. Mixture-of-Transformers reasoner and generator with shared attention. Source paper, PDF p. 11. Figure notice.
ℹ️ 詳細情報
Cosmos 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3。
Cosmos 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、32B、20、3。
Lance: Unified Image and Video Understanding, Generation, and Editing
Lanceの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Lance Team
Released: 2026-05-18
Figure 6. Dual-expert sequence modeling for understanding and visual generation. Source paper, PDF p. 9. Figure notice.
ℹ️ 詳細情報
Lanceの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3B。
Lanceの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 128。
ZAYA1-VL: Vision-Specialized Compressed Convolutional Attention
ZAYA1-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Zyphra
Released: 2026-05-08
Figure 2. Visual routing, compressed convolutional attention, and a hybrid language backbone. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
ZAYA1-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、5-。
ZAYA1-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 140B、2.0。
Falcon Perception: Early-Fusion Grounding, Segmentation, and OCR
Falcon Perceptionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Technology Innovation Institute
Released: 2026-05-03
Figure 1. Early-fusion perception Transformer with grounding, geometry, and segmentation pathways. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
Falcon Perceptionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Falcon Perceptionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 600M、300M、28、3。
GLM-5V-Turbo: Native Multimodal Agency
GLM-5V-Turboの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
GLM-V Team
Released: 2026-04-29
Figure 2. Multimodal multi-token prediction with image placeholders and shared Transformer blocks. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
GLM-5V-Turboの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
GLM-5V-Turboの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.。
PLaMo 2.1-VL: Lightweight Japanese Vision-Language Modeling
PLaMo 2.1-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.1-、2B、8B。
Tommi Kerola et al., Preferred Networks
Released: 2026-04-21
Architecture figure: The PLaMo 2.1-VL paper contains application and data figures, but no model architecture diagram.
ℹ️ 詳細情報
PLaMo 2.1-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.1-。
PLaMo 2.1-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
EXAONE 4.5: Native Multimodal Pretraining for Documents
EXAONE 4.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.5。
Eunbi Choi et al., LG AI Research
Released: 2026-04-09
Figure 1. Native-resolution vision encoding, projection, language decoding, and multi-token prediction. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
EXAONE 4.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.5、4.0。
EXAONE 4.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 256K。
BidirLM and BidirLM-Omni: Causal Decoders as Multimodal Encoders
BidirLM and BidirLM-Omniの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
BidirLM Team
Released: 2026-04-02
Figure 10. Specialist-backbone merging with frozen modality projection heads. Source paper, PDF p. 25. Figure notice.
ℹ️ 詳細情報
BidirLM and BidirLM-Omniの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
BidirLM and BidirLM-Omniの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Gemma 4: Open-Weight Native Multimodal Models
Gemma 4の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4。
Gemma Team, Google DeepMind
Released: 2026-03-31
Figure 2. Aspect-preserving image resizing, patch pooling, and soft-token production. Source paper, PDF p. 16. Figure notice.
ℹ️ 詳細情報
Gemma 4の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4、31B、26B、4B、12B、128K、256K。
Gemma 4の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Gemma 4の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Penguin-VL: Efficient VLMs with LLM-Based Vision Encoders
Penguin-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Boqiang Zhang, Lei Ke, Ruihan Yang, Qi Gao, Tianyuan Qu, Rossell Chen, Dong Yu, Leoweiliang
Released: 2026-03-06
Figure 3. An LLM-initialized vision encoder with priority-aware video-token compression. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
Penguin-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 0.6B、2B、8B。
Penguin-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Penguin-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Phi-4-Reasoning-Vision: Compact Multimodal Reasoning
Phi-4-Reasoning-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4-、15B。
Jyoti Aneja, Michael Harrison, Neel Joshi, Tyler LaBonte, John Langford, Eduardo Salinas
Released: 2026-03-04
Figure 3. SigLIP2 vision encoding, cross-modal projection, mid-fusion, and language reasoning. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
Phi-4-Reasoning-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4-、15B、2、3,600。
Phi-4-Reasoning-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連コード: <think>、<nothink>。 値: 240。
Phi-4-Reasoning-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
V-SONAR and V-LCM: Vision-Language Modeling in Concept Space
V-SONAR and V-LCMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Yifu Qiu, Paul-Ambroise Duquenne, Holger Schwenk
Released: 2026-03-01
Figure 1. Visual-semantic alignment and concept-space prediction. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
V-SONAR and V-LCMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
V-SONAR and V-LCMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
V-SONAR and V-LCMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 80、1K。
Qwen3.5: Native Multimodal Hybrid-Attention Models
Qwen3.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、6。
Qwen Team
Released: 2026-02-16
Architecture figure: Qwen3.5 has no public technical paper containing an architecture figure.
ℹ️ 詳細情報
Qwen3.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、397B、17B。
Qwen3.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Qwen3.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 6、5。
Youtu-VL: Unified Autoregressive Supervision for Dense Vision
Youtu-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Tencent Youtu Lab
Released: 2026-01-27
Figure 3. Unified visual-text autoregressive supervision and dense-output decoding. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
Youtu-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Youtu-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Kimi K2.5 and K2.6: Native Multimodal Agentic MoE
Kimi K2.5 and K2.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、6。
Kimi Team, Moonshot AI
Released: 2026-01-27
Figure 10. Agentic reinforcement-learning environments, rollout management, and training services. Source paper, PDF p. 23. Figure notice.
ℹ️ 詳細情報
Kimi K2.5 and K2.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、1T、32B、61、384。
Kimi K2.5 and K2.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 20、6、5、256K。
Step3-VL-10B: Language-Aligned Perception with 16× Token Compression
Step3-VL-10Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 10B、1.8B、8B。
StepFun
Released: 2026-01-14
Architecture figure: The Step3-VL-10B report contains performance and RL figures, but no architecture diagram.
ℹ️ 詳細情報
Step3-VL-10Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 10B、2025、1.8B、2、16、8B。
Step3-VL-10Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 728×728、504×504。
ERNIE 5.0: Unified Autoregressive Omnimodal Mixture of Experts
ERNIE 5.0の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5.0、2.4T。
Baidu ERNIE Team
Released: 2025-11-13
Figure 2. Unified image understanding, image generation, and video generation objectives. Source paper, PDF p. 6. Figure notice.
ℹ️ 詳細情報
ERNIE 5.0の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5.0、3、2.4T。
ERNIE 5.0の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5.0、13、2025、2026、4。
DeepSeek-OCR: Visual Context Compression through DeepEncoder
DeepSeek-OCRの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
DeepSeek-AI
Released: 2025-10-20
Figure 3. A SAM-CLIP DeepEncoder connected to a sparse language decoder. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
DeepSeek-OCRの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3B、570M、64、800。
DeepSeek-OCRの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
DeepSeek-OCRの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連参照: https://arxiv.org/abs/2601.20552、https://github.com/deepseek-ai/DeepSeek-OCR-2。 値: 2、27、2026。
PaddleOCR-VL: Ultra-Compact Multilingual Document Parsing
PaddleOCR-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.5-0.3B、0.9B。
PaddleOCR Team, Baidu
Released: 2025-10-16
Figure 2. Document layout analysis, compact VLM inference, instructions, and structured output. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
PaddleOCR-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 0.9B、4.5-0.3B、109。
PaddleOCR-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5、2026、1.6。
Qwen3-VL: DeepStack Vision-Language Models
Qwen3-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Qwen Team
Released: 2025-09-22
Figure 1. Vision encoding, DeepStack injection, and dense or mixture-of-experts decoding. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
Qwen3-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2B、4B、8B、32B、30B、235B、256K。
Qwen3-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Qwen3-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Step3: Model-System Co-Design for Cost-Effective Multimodal Intelligence
Step3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 321B、38B。
StepFun
Released: 2025-07-25
Figure 6. Attention-FFN disaggregation across attention and expert instances. Source paper, PDF p. 11. Figure notice.
ℹ️ 詳細情報
Step3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 321B、38B。
Step3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
GLM-4.1V-Thinking: General-Purpose Multimodal Reasoning through Curriculum-Sampled RL
GLM-4.1V-Thinkingの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.、9B。
GLM-V Team, Zhipu AI and Tsinghua University
Released: 2025-07-01
Figure 2. Native-resolution vision encoding, projection, decoding, and timestamped video tokens. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
GLM-4.1V-Thinkingの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.、9B、4-9B、0414。
GLM-4.1V-Thinkingの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連コード: <think>、<answer>。 値: 32K。
GLM-4.1V-Thinkingの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 220M、4.。
ERNIE 4.5-VL: Heterogeneous Modality Mixture-of-Experts
ERNIE 4.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.5-。
Baidu ERNIE Team
Released: 2025-06-30
Architecture figure: ERNIE 4.5-VL has no public paper containing an extractable architecture figure.
ℹ️ 詳細情報
ERNIE 4.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.5-、424B、47B、28B、3B、4.5。
ERNIE 4.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
MiMo-VL: Multimodal Pretraining with Mixed On-Policy Reinforcement Learning
MiMo-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.4T。
Xiaomi MiMo Team
Released: 2025-06-04
Figure 2. Native-resolution vision encoding, projection, and language decoding. Source paper, PDF p. 6. Figure notice.
ℹ️ 詳細情報
MiMo-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B。
MiMo-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.4T。
BAGEL: A Mixture-of-Transformer-Experts for Unified Understanding and Generation
BAGELの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, Guang Shi, Haoqi Fan
Released: 2025-05-20
Figure 2. Shared self-attention with understanding and generation Transformer experts. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
BAGELの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14B、7B、5-、2。
BAGELの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
BAGELの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 400M、500M、1.6B、100M、45M、20M、500K。
Seed1.5-VL: Sparse-MoE Multimodal Understanding and Agentic Reasoning
Seed1.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5-、532M、20B。
ByteDance Seed Team
Released: 2025-05-11
Figure 1. Native-resolution vision encoding, adaptation, sparse MoE decoding, and timestamped video. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
Seed1.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5-、20B。
Seed1.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
InternVL3 and InternVL3.5: Native Multimodal Pretraining and Adaptive Resolution
InternVL3 and InternVL3.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5。
InternVL Team, OpenGVLab
Released: 2025-04-11
Architecture figure: The InternVL3 paper contains evaluation figures, but no definitive model architecture diagram.
ℹ️ 詳細情報
InternVL3 and InternVL3.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
InternVL3 and InternVL3.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 26、2025、5。
Kimi-VL: Native-Resolution Vision with a Sparse MoE Decoder
Kimi-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.8B、16B、128K。
Kimi Team
Released: 2025-04-10
Figure 3. MoonViT, multimodal projection, and a sparse mixture-of-experts decoder. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
Kimi-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 16B、2.8B、2×2。
Kimi-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5.2T、4.4T、2T、0.1T、2.3T、8K、128K、2506。
Kimi-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Llama 4 Scout and Maverick: Native Multimodal Mixture-of-Experts Models
Llama 4 Scout and Maverickの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4、17B。
Meta
Released: 2025-04-05
Architecture figure: The official Llama 4 model card describes the architecture in text and tables only.
ℹ️ 詳細情報
Llama 4 Scout and Maverickの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4、109B、16、400B、128、17B、10M、1M。
Llama 4 Scout and Maverickの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 40T、22T、4。
Qwen2.5-Omni: Streaming Multimodal Perception and Speech Generation
Qwen2.5-Omniの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5-。
Qwen Team
Released: 2025-03-26
Figure 2. Thinker-Talker architecture for multimodal perception and streaming speech. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
Qwen2.5-Omniの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5-。
Qwen2.5-Omniの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Gemma 3: Long-Context Multimodality with Efficient Interleaved Attention
Gemma 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3、128K。
Gemma Team
Released: 2025-03-12
Architecture figure: The Gemma 3 report contains examples and analysis charts, but no architecture diagram.
ℹ️ 詳細情報
Gemma 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3、1B、4B、12B、27B、400M、896×896、256。
Gemma 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3、2、1,024-、32K、128K。
Gemma 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2T、4T、12T、14T。
Aya Vision: Multilingual Multimodality through Cross-Modal Model Merging
Aya Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Cohere Labs
Released: 2025-03-04
Architecture figure: The Aya Vision paper contains data and evaluation figures, but no model architecture diagram.
ℹ️ 詳細情報
Aya Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、32B、23。
Aya Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Phi-4-multimodal: Text, Vision, and Speech through Mixture-of-LoRAs
Phi-4-multimodalの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4-。
Microsoft Phi Team
Released: 2025-03-03
Figure 1. Vision and audio encoders, projectors, and modality-specific LoRA routes. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
Phi-4-multimodalの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4-、5.6-、460。
Phi-4-multimodalの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 200K、128K。
Phi-4-multimodalの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3.5、4。
SigLIP 2: Multilingual Vision-Language Encoders with Native-Aspect-Ratio Support
SigLIP 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。
Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muhammad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, et al.
Released: 2025-02-20
Figure 1. Combined contrastive, captioning, masked-prediction, and self-distillation objectives. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
SigLIP 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。
SigLIP 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
SigLIP 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 10、12、109、90、2。
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
EVEv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
EVEv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 3. EVEv2.0 architecture: lossless patch embeddings and text tokens enter a unified decoder-only VLM whose attention, feed-forward, and normalization layers use modality-specific weights. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
EVEv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、10M。
Qwen2.5-VL: Enhanced Vision-Language Capabilities in the Qwen Series
Qwen2.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5-。
Qwen2.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. Qwen2.5-VL framework: a vision encoder processes native-resolution images and dynamic-FPS video into variable-length tokens for a Qwen2.5 language-model decoder with multimodal rotary position encoding. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
Qwen2.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5-、3B、7B、72B、18。
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
VideoLLaMA 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
VideoLLaMA 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 3. VideoLLaMA 3 pipeline with any-resolution vision tokenization and a differential frame pruner that compresses redundant video tokens before language-model processing. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
VideoLLaMA 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、1、2、3、4、1-。
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
UI-TARSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 10。
UI-TARSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 4. UI-TARS architecture and capability overview, connecting visual GUI observations and interaction histories to perception, action, system-level reasoning, and learning from prior experience. Source paper, PDF p. 14. Figure notice.
ℹ️ 詳細情報
UI-TARSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1、2、3、4、2-、7B、72B。
MiniMax-01: Scaling Foundation Models with Lightning Attention
MiniMax-01の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 01、3.5-、4。
MiniMax-01の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 3. MiniMax-Text-01 backbone architecture, interleaving Lightning Attention and softmax-attention transformer blocks with routed mixture-of-experts feed-forward layers. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
MiniMax-01の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 01、456、45.9、32、2、4。 MiniMax-01の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 01、694、100、2。
MiniCPM-o-2.6: A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming
MiniCPM-o-2.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.6、8B。
MiniCPM-o-2.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Official architecture diagram. MiniCPM-o 2.6 end-to-end omni-modal streaming architecture, using time-division multiplexing to combine visual, audio, and query streams in a shared backbone with streaming speech decoding. Primary source. Figure notice.
ℹ️ 詳細情報
MiniCPM-o-2.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.6、400M、300M、200M、5-7B。
MiniCPM-o-2.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連参照: https://arxiv.org/abs/2604.27393、https://github.com/OpenBMB/MiniCPM-V。 値: 4.5、2026。
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
Eagle 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。
Eagle 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 11. Eagle 2's tiled mixture of vision encoders, combining SigLIP and ConvNeXt features through dynamic image splitting, feature concatenation, pixel shuffle, and an MLP connector to the LLM. Source paper, PDF p. 7. Figure notice.
ℹ️ 詳細情報
Eagle 2 adopts a “diversity first, then quality” data strategy, beginning with a large, diverse pool of over 180 data sources, followed by rigorous filtering and selection. The architecture uses a tiled mixture of vision encoders (MoVE), specifically SigLIP and ConvNeXt-XXLarge, with image tiling to handle high resolutions. Each image tile is encoded by channel-concatenated MoVE. The vision encoder outputs are concatenated and aligned with the LLM (Qwen2.5) via a simple MLP connector. A three-stage training recipe is used: Stage 1 trains the connector to align modalities; Stage 1.5 trains the full model on a large, diverse dataset; and Stage 2 fine-tunes on a high-quality instruction-tuning dataset. Crucially, all available visual instruction data is used in Stage 1.5, not just captioning/knowledge data. Balanced data packing addresses limitations in existing open-source frameworks. The core contribution is the detailed data strategy. This involves: (1) Data Collection: Building a highly diverse data pool (180+ sources) through both passive gathering (monitoring arXiv and Hugging Face) and proactive searching (addressing “bucket effect” via error analysis). (2) Data Filtering: Removing low-quality samples based on criteria like mismatched question-answer pairs, irrelevant image-question pairs, repeated text, and numeric formatting issues. (3) Data Selection: Choosing optimal subsets based on data source diversity, distribution, and K-means clustering on SSCD image embeddings to ensure balance across types (especially useful for chart data, etc.). (4) Data Augmentation: Mining information from input images through techniques like Chain-of-Thought (CoT) explanation generation, rule-based QA generation, and expanding short answers into longer ones. (5) Data Formating: remove unnecessary decorations. Training uses a three-stage approach: Eagle 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1。 Eagle 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5、21.6M。 Eagle 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、4.6M。
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
Sa2VAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。
Sa2VAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. Sa2VA model: text, prompts, images, and videos are encoded for an LLM, whose segmentation token is combined with SAM 2 features to decode image or video masks. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
Sa2VAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、91、93、72,000、2,000、1.5、665K、17K、22K、214K、100K、3.5K、0.6K、1.7K、37K。
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
VideoChat-Flashの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
VideoChat-Flashの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 3. VideoChat-Flash framework with hierarchical video-token compression: shared encoders and connectors first compress clips, then the LLM performs video-level compression for long-context inference. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
VideoChat-Flashの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、224、1、0.5M、2、3.5M、2.5M、3、1.1M、1.7M、0.7M、4、448、25、15、114,228、3,444,849。
OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference
OmniVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 729、81、5-0.5B、400M。
OmniVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. OmniVLM architecture: a vision transformer feeds a reshape-based projector that compresses image tokens before they join text tokens in the Qwen2.5-0.5B-Instruct language model. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
OmniVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 91、729、93、81、1、400M、5-0.5B、2、3、81-、9.、1.。
Apollo: An Exploration of Video Understanding in Large Multimodal Models
Apolloの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3B、7B。
Apolloの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 8. Apollo architecture: image and video encoders process N-frame clips, interpolated features are concatenated channel-wise, and a Perceiver resampler produces a fixed token set for the language model. Source paper, PDF p. 21. Figure notice.
ℹ️ 詳細情報
Apolloの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、1.5B、3B、7B、32、3-、8-32。 or clips is sufficient for efficient token integration. Training Stages is also disscussed, concluding that progressively unfreezing the different components in different stages leads to superior model training dynamics. Finally, training the Video Encoder is discussed. The paper concludes that Finetuning video encoders on only video data further improves overall performance, Apolloの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 Apolloの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
DeepSeek-VL2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
DeepSeek-VL2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. DeepSeek-VL2 architecture: dynamic image tiling feeds a vision encoder and vision-language adapter whose image tokens join text tokens in a DeepSeek-MoE language model. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
DeepSeek-VL2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1、2、3、384、9、729、1152、196、14、210、1.0B、2.8B、4.5B、1.2M、70、30、12M、800B。
Maya: An Instruction Finetuned Multilingual Multimodal Model
Mayaの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Mayaの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 7. Maya's LLaVA-derived architecture, projecting multilingual SigLIP vision features into the embedding space of a multilingual language model for instruction following. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
Mayaの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連コード: Zv = g(Xv)、W、Hv。 値: 1.5、23、8B、2-、35B、7,531、3、150K、10、2、558,000、7B、13B。
InternVL 2.5: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
InternVL 2.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.5、2.0、3.5-。
InternVL 2.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. InternVL 2.5's ViT-MLP-LLM architecture with pixel-unshuffle visual-token compression. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
InternVL 2.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連コード: nmax、r。 値: 2.5、6B、300M、2-、1024、256、2.0、1、2、3。
PaliGemma 2: A Family of Versatile VLMs for Transfer
PaliGemma 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、3B、10B、28B。
PaliGemma 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. PaliGemma 2 processes variable-resolution images with SigLIP, a linear projector, and Gemma 2. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
PaliGemma 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、2B、9B、27B、3B、10B、28B、1、50、10、3。
ShowUI: Vision-Language-Action Modeling for GUI Agents
ShowUIの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2B。
Kevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang, Shiwei Wu, Zechen Bai, Weixian Lei, Lijuan Wang, Mike Zheng Shou
Released: 2024-11-26
Figure 3. UI-guided visual-token selection and interleaved action history. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
ShowUIの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2B、33、1.4。
ShowUIの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 256K。
ShowUIの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
SmolVLM: A Small, Efficient, and Open-Source Vision-Language Model
SmolVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2B、2.0。
SmolVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Official architecture diagram. SmolVLM's SigLIP vision encoder, aggressive pixel-shuffle compression, and SmolLM2 language backbone. Primary source. Figure notice.
ℹ️ 詳細情報
SmolVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3.1、8B、1.7B、9、1.。
SmolVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連参照: https://huggingface.co/blog/smolvlm2。 値: 20、2025、256M、500M、2.2B。
AIMv2: Multimodal Autoregressive Pre-training of Large Vision Encoders
AIMv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
AIMv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. AIMv2's prefix-attention vision encoder and joint autoregressive multimodal decoder. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
AIMv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2B、300、3。
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
LLaVA-CoTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
LLaVA-CoTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 4. LLaVA-CoT's Best-of-N, stage-wise beam-search, and stage-wise retracing inference procedures. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
LLaVA-CoTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3.2-。
LLM2CLIP: Powerful Language Model Unlocks Richer Visual Representation
LLM2CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
LLM2CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. LLM2CLIP fine-tunes an LLM for caption discrimination before using it to train stronger CLIP representations. Source paper, PDF p. 1. Figure notice.
ℹ️ 詳細情報
LLM2CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1、3、8B、2。 LLM2CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 15M、3M、103、12M、1B。
Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Tarsier2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5。
Tarsier2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Architecture figure: The paper contains capability, dataset, DPO-construction, and benchmark figures but inherits Qwen2-VL and provides no architecture overview; a training diagram would be misleading.
ℹ️ 詳細情報
Tarsier2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 11、40、150K、15、585K。
Janus and Janus-Pro: Decoupled Visual Understanding and Generation
Janus and Janus-Proの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
DeepSeek-AI
Released: 2024-10-17
Figure 3. Janus-Pro decouples visual understanding and generation encoders around a shared autoregressive transformer. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
Janus and Janus-Proの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Janus and Janus-Proの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2025、1.5B、7B、90M、72M。
ARIA: An Open Multimodal Native Mixture-of-Experts Model
ARIAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4-。
ARIAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Architecture figure: ARIA’s visual encoder, projection layer, and fine-grained MoE are described in prose and a configuration table; its only numbered model figure visualizes expert specialization, not architecture.
ℹ️ 詳細情報
ARIAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3.5B、3.9B、24.9B、66、2、6、438M、4-、1、6.4T、8K、400B、3、64K、4、20B。 ARIAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Emu3: Next-Token Prediction across Text, Image, and Video
Emu3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B。
Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, Yingli Zhao, Yulong Ao, Xuebin Min, Tao Li, Boya Wu, et al.
Released: 2024-09-27
Figure 1. A single next-token objective across text, image, and video tokens. Source paper, PDF p. 1. Figure notice.
ℹ️ 詳細情報
Emu3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Emu3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、0.2、0.8。
Emu3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Molmo and PixMo: Open Weights, Open Data, and Grounded Pointing
Molmo and PixMoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, et al.
Released: 2024-09-25
Figure 2. Vision encoding, multimodal connection, language decoding, and grounded pointing. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
Molmo and PixMoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14、1B、7B、72B。
Molmo and PixMoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Molmo and PixMoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Llama 3.2-Vision: Enhanced Multimodal Capabilities Built on Llama 3
Llama 3.2-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3.2-、3、11B、90B。
Llama 3.2-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Architecture figure: The official source describes the image encoder and cross-attention adapter in prose; benchmark tables and marketing artwork are not legitimate substitutes.
ℹ️ 詳細情報
Llama 3.2-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連コード: <|image|>、brave_search、wolfram_alpha。 値: 3.2-、3、6B、2023、3.2。
NVLM: Open Frontier-Class Multimodal LLMs
NVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.0、1-。
NVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 3. NVLM-X, NVLM-H, and NVLM-D share a dynamic-high-resolution visual pathway but integrate vision differently. Source paper, PDF p. 9. Figure notice.
ℹ️ 詳細情報
NVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連コード: <tile_1>。 値: 6B、5、6、1024、256、72B、2-、34B、1-、115M、40、40-。
Pixtral 12B: A Cutting-Edge Open Multimodal Language Model
Pixtral 12Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 12B、12-。
Pixtral 12Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 3. Pixtral combines a variable-resolution vision encoder with a 128K-context multimodal decoder. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
Pixtral 12Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 12B、128K、91、93、12-。
Pixtral 12Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連参照: https://mistral.ai/news/pixtral-large、https://mistral.ai/news/mistral-small-4/。 値: 2024、4、2026。
VILA-U: Fully Autoregressive Visual Understanding and Generation
VILA-Uの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
VILA-U Team
Released: 2024-09-06
Figure 1. A shared visual tokenizer, autoregressive model, and modality decoders. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
VILA-Uの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
VILA-Uの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Qwen2-VL: A Powerful Open-Source Vision-Language Model for Image and Video Understanding
Qwen2-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Bai, Jinze and Bai, Shuai and Yang, Shusheng and Wang, Shijie and Tan, Sinan and Wang, Peng and Lin, Junyang and Zhou, Chang and Zhou, Jingren
Figure 3. Qwen2-VL's M-RoPE decomposes multimodal position encoding into temporal, height, and width components. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
Qwen2-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 600M、20。
EAGLE: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
EAGLEの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Min Shi, Fuxiao Liu, Shihao Wang, Shijia Liao, Subhashree Radhakrishnan, De-An Huang, Hongxu Yin, Karan Sapra, Yaser Yacoob, Humphrey Shi, Bryan Catanzaro, Andrew Tao, Jan Kautz, Zhiding Yu, Guilin Liu
Figure 2. EAGLE explores mixtures of vision experts and alternative fusion strategies for multimodal language models. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
EAGLEの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Show-o: Autoregressive Language and Discrete-Diffusion Vision in One Transformer
Show-oの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Show Lab
Released: 2024-08-22
Figure 2. Causal text attention and full-attention discrete image diffusion. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
Show-oの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Show-oの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Idefics3-8B: Building and Better Understanding Vision-Language Models
Idefics3-8Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B。
Hugo Laurençon, Andrés Marafioti, Victor Sanh, Léo Tronchon
Figure 1. Idefics3 maps vision-encoder features into interleaved visual tokens consumed by an autoregressive language model. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
Idefics3-8Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、3.1、1.5、4、169、364、1820、13.7-。
Transfusion: Next-Token Text Prediction and Continuous Image Diffusion
Transfusionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Meta FAIR
Released: 2024-08-20
Figure 1. A shared Transformer with autoregressive text and continuous image diffusion. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
Transfusionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Transfusionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、16。
mPLUG-Owl3: Hyper-Attention for Long Image Sequences
mPLUG-Owl3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
mPLUG-Owl Team
Released: 2024-08-09
Figure 2. Vision encoding and Hyper-Attention blocks inside the language model. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
mPLUG-Owl3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
mPLUG-Owl3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
VITA: Towards Open-Source Interactive Omni Multimodal LLM
VITAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
VITAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. VITA unifies text, audio, image, and video inputs with state tokens, an LLM, and speech output. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
VITAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 300M、256、24-、25、2。
LLaVA-OneVision: Easy Visual Task Transfer
LLaVA-OneVisionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, Chunyuan Li
Figure 1. LLaVA-OneVision extends the minimal LLaVA vision-encoder, projector, and LLM architecture to multiple visual signals. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
LLaVA-OneVisionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、2-。
LLaVA-OneVisionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連参照: https://arxiv.org/abs/2509.23661、https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-1.5。 値: 28、2025、1.5、85M、26M、64B、16,000。
VILA²: VILA Augmented VILA
VILA²の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Yunhao Fang, Ligeng Zhu, Yao Lu, Yan Wang, Pavlo Molchanov, Jang Hyun Cho, Marco Pavone, Song Han, Hongxu Yin
Figure 1. VILA² improves itself through generic model-in-the-loop recaptioning and specialist-model augmentation. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
VILA²の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、8B、34B、6B。
INF-LLaVA: High-Resolution Image Perception for Multimodal Large Language Models
INF-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Yiwei Ma, Zhibin Wang, Xiaoshuai Sun, Weihuang Lin, Qiang Zhou, Jiayi Ji, Rongrong Ji
Figure 2. INF-LLaVA combines dual-perspective cropping, CLIP encoding, feature recombination, enhancement, and language reasoning. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
INF-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
SlowFast-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Mingze Xu, Mingfei Gao, Zhe Gan, Hong-You Chen, Zhengfeng Lai, Haiming Gang, Kai Kang, Afshin Dehghan
Figure 2. SlowFast-LLaVA combines low-frame-rate spatial detail with high-frame-rate motion features. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
SlowFast-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8、64。
EVLM: An Efficient Vision-Language Model for Visual Understanding
EVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Kaibing Chen, Dong Shen, Hanwen Zhong, Huasong Zhong, Kui Xia, Di Xu, Wei Yuan, Yifei Hu, Bin Wen, Tianke Zhang, Changyi Liu, Dewen Fan, Huihui Xiao, Jiahong Wu, Fan Yang, Size Li, Di Zhang
Figure 2. EVLM injects hierarchical EVA2-CLIP features into the language model through gated cross-attention layers. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
EVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.4B、8、40、16、14B、1.0。
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
InternLM-XComposer-2.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.5、7B。
Pan Zhang, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Rui Qian, Lin Chen, Qipeng Guo, Haodong Duan, Bin Wang, Linke Ouyang, Songyang Zhang, Wenwei Zhang, Yining Li, Yang Gao, Peng Sun, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Hang Yan, Conghui He, Xingcheng Zhang, Kai Chen, Jifeng Dai, Yu Qiao, Dahua Lin, Jiaqi Wang
Figure 5. InternLM-XComposer-2.5's framework supports text, single and multiple images, and video inputs. Source paper, PDF p. 6. Figure notice.
ℹ️ 詳細情報
InternLM-XComposer-2.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.5、2、2-、14、7B。
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
OMG-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Tao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan, Shengqiong Wu, Shunping Ji, Chen Change Loy, Shuicheng Yan
Figure 3. OMG-LLaVA connects OMG-Seg visual tokens and prompts to an LLM that can decode segmentation outputs. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
OMG-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Cambrian-1: Vision-Centric Multimodal LLMs
Cambrian-1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1、8B、13B、34B。
Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Ziteng Wang, Rob Fergus, Yann LeCun, Saining Xie
Released: 2024-06-24
Figure 8. Spatial Vision Aggregation across multiple visual encoders and decoder layers. Source paper, PDF p. 13. Figure notice.
ℹ️ 詳細情報
Cambrian-1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1、3、1.5、2-。
Cambrian-1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.5M、7M。
Cambrian-1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7M、10M。
EVE: Unveiling Encoder-Free Vision-Language Models
EVEの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 35M。
EVEの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. EVE combines patch embedding, a decoder-only backbone, patch alignment, and next-word prediction without a separate vision encoder. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
EVEの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連コード: <CLS>、<SPL>。 値: 7B、14、16M、33M、665K、1。
EVEの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 33M、1.5、665K。
Ovis: Structural Visual-Text Embedding Alignment
Ovisの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Alibaba International Digital Commerce
Released: 2024-06-14
Figure 3. Visual-token probability distributions and structural visual embedding lookup. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
Ovisの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、1.6。
Ovisの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 26、2025、5、1B、34B、15、2。
Parrot: Multilingual Visual Instruction Tuning
Parrotの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Hai-Long Sun, Da-Wei Zhou, Yang Li, Shiyin Lu, Chao Yi, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, De-Chuan Zhan, Han-Jia Ye
Figure 5. PARROT aligns multilingual visual features through a projector, multilingual mixture of experts, and language model. Source paper, PDF p. 6. Figure notice.
ℹ️ 詳細情報
Parrotの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14、5-。
ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models
ConvLLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Chunjiang Ge, Sijie Cheng, Ziming Wang, Jiale Yuan, Yuan Gao, Jun Song, Shiji Song, Gao Huang, Bo Zheng
Figure 1. ConvLLaVA uses a hierarchical ConvNeXt vision encoder to compress visual tokens between stages. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
ConvLLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 576。
Phi-3-Vision and Phi-3.5-Vision: Compact Long-Context Multimodal Reasoning
Phi-3-Vision and Phi-3.5-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3-、4.2B、3.5-。
Microsoft Phi Team
Released: 2024-05-21
Architecture figure: The cited Phi-3 report contains no Phi-3-Vision architecture diagram.
ℹ️ 詳細情報
Phi-3-Vision and Phi-3.5-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3-、3、128K、4.2B。
Phi-3-Vision and Phi-3.5-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3-、21、2024、3.5-、22。
CogVLM2: Enhanced Vision-Language Models for Image and Video Understanding
CogVLM2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Wenyi Hong, Weihan Wang, Ming Ding, Wenmeng Yu, Qingsong Lv, Yan Wang, Yean Cheng, Shiyu Huang, Junhui Ji, Zhao Xue, Lei Zhao, Zhuoyi Yang, Xiaotao Gu, Xiaohan Zhang, Guanyu Feng, Da Yin, Zihan Wang, Ji Qi, Xixuan Song, Peng Zhang, Debing Liu, Bin Xu, Juanzi Li, Yuxiao Dong, Jie Tang
Figure 2. CogVLM2 processes high-resolution images and video frames through a ViT encoder, adapter, and visual-language decoder. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
CogVLM2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Chameleon: Mixed-Modal Early-Fusion Foundation Models
Chameleonの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Chameleon Team
Released: 2024-05-16
Figure 1. Mixed-modal early-fusion tokenization and autoregressive generation. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
Chameleonの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、34B。
Chameleonの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.4。
Chameleonの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
PaliGemma: A Versatile and Transferable 3B Vision-Language Model
PaliGemmaの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2B、40。
Lucas Beyer, Andreas Steiner, André Susano Pinto, Alexander Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, Thomas Unterthiner, Daniel Keysers, Skanda Koppula, Fangyu Liu, Adam Grycner, Alexey Gritsenko, Neil Houlsby, Manoj Kumar, Keran Rong, Julian Eisenschlos, Rishabh Kabra, Matthias Bauer, Matko Bošnjak, Xi Chen, Matthias Minderer, Paul Voigtlaender, Ioana Bica, Ivana Balazevic, Joan Puigcerver, Pinelopi Papalampidi, Olivier Henaff, Xi Xiong, Radu Soricut, Jeremiah Harmsen, Xiaohua Zhai
Figure 1. PaliGemma connects a SigLIP image encoder to a Gemma autoregressive decoder language model. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
PaliGemmaの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3-、2B、30。
xGen-MM (BLIP-3): An Open-Source Framework for Building Powerful and Responsible Large Multimodal Models
xGen-MM (BLIP-3)の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3。
Le Xue, Manli Shu, Anas Awadalla, Jun Wang, An Yan, Senthil Purushwalkam, Honglu Zhou, Viraj Prabhu, Yutong Dai, Michael S Ryoo, Shrikant Kendre, Jieyu Zhang, Can Qin, Shu Zhang, Chia-Chih Chen, Ning Yu, Juntao Tan, Tulika Manoj Awalgaonkar, Shelby Heinecke, Huan Wang, Yejin Choi, Ludwig Schmidt, Zeyuan Chen, Silvio Savarese, Juan Carlos Niebles, Caiming Xiong, Ran Xu
Figure 2. BLIP-3 feeds interleaved visual and text tokens through a scalable vision-token sampler into a pretrained language model. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
xGen-MM (BLIP-3)の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3、2。
MANTIS: Mastering Multi-Image Understanding Through Interleaved Instruction Tuning
MANTISの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Dongfu Jiang, Xuan He, Huaye Zeng, Cong Wei, Max Ku, Qian Liu, Wenhu Chen
Architecture figure: The paper contains capability examples, dataset statistics and illustrations, and case studies, none of which is a model architecture or system diagram.
ℹ️ 詳細情報
MANTISの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 721K、3、8B。
Moondream-next: Compact Vision-Language Model with Enhanced Capabilities
Moondream-nextの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.9B。
Architecture figure: The rolling model card and official repository provide implementation details but no legitimate architecture or system figure.
ℹ️ 詳細情報
Moondream-nextの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.9B。
Idefics2
Idefics2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、7B。
Idefics2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. Idefics2 maps vision-encoder features into interleaved visual tokens consumed by an autoregressive language model. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
Idefics2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、7B、2.5。
InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
InternLM-XComposer2-4KHDの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 336、4K。
Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Songyang Zhang, Haodong Duan, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Zhe Chen, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Kai Chen, Conghui He, Xingcheng Zhang, Jifeng Dai, Yu Qiao, Dahua Lin, Jiaqi Wang
Figure 4. InternLM-XComposer2-4KHD dynamically partitions high-resolution images into local patches while retaining a global thumbnail. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
InternLM-XComposer2-4KHDの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4K、336。
MM1: Methods, Analysis, and Insights from Multimodal Pre-training
MM1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Floris Weers, Anton Belyi, Haotian Zhang, et al.
Released: 2024-03-14
Architecture figure: MM1 presents an architecture design space and ablations, but no definitive final-model diagram.
ℹ️ 詳細情報
MM1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3B、30B。
MM1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
MM1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
DeepSeek-VL: Towards Real-World Vision-Language Understanding
DeepSeek-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, Yaofeng Sun, Chengqi Deng, Hanwei Xu, Zhenda Xie, Chong Ruan
Figure 3. DeepSeek-VL trains its hybrid vision encoder, vision-language adaptor, and language model across three progressive stages. Source paper, PDF p. 12. Figure notice.
ℹ️ 詳細情報
DeepSeek-VL: Employs a hybrid vision encoder architecture, fusing a SigLIP-L encoder for semantic understanding with a SAM-B encoder for high-resolution detail extraction. This allows for efficient processing of 1024x1024 images while capturing both global and fine-grained visual features. A two-layer hybrid MLP adapter then integrates these features with the DeepSeek LLM backbone. The model is pre-trained on a diverse dataset encompassing web screenshots, PDFs, OCR, charts, and knowledge-based content from sources like Common Crawl, Web Code, E-books, and arXiv articles. This pretraining is further refined using a curated instruction-tuning dataset based on real user scenarios and categorized into a comprehensive taxonomy covering recognition, conversion, analysis, reasoning, evaluation, and safety tasks. By combining this diverse data with its unique architecture and fusion strategies, DeepSeek-VL aims to deliver robust performance across a wide range of real-world vision-language applications.
AnyGPT: Unified Any-to-Any Multimodal Modeling with Discrete Tokens
AnyGPTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
OpenMOSS
Released: 2024-02-19
Figure 1. Unified discrete sequence modeling across speech, text, image, and music. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
AnyGPTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
AnyGPTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 108,000-。
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
SPHINX-Xの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
SPHINX-Xの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 3. SPHINX-X combines mixed visual experts, padded-tile skip tokens, high-resolution partitioning, and unified training. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
SPHINX-Xの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
LLaVA 1.6: LLaVA-NeXT Improved reasoning, OCR, and world knowledge
LLaVA 1.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5。
LLaVA 1.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Official architecture diagram. LLaVA-NeXT's AnyRes scheme partitions high-resolution images into a configurable grid of local views. Primary source. Figure notice.
ℹ️ 詳細情報
LLaVA 1.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5、1、32。
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
MiniCPM-Vの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.5、3。
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, Qianyu Chen, Huarong Zhou, Zhensheng Zou, Haoye Zhang, Shengding Hu, Zhi Zheng, Jie Zhou, Jie Cai, Xu Han, Guoyang Zeng, Dahai Li, Zhiyuan Liu, Maosong Sun
Figure 3. MiniCPM-V combines a visual encoder, shared compression layer, language model, and adaptive high-resolution encoding. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
MiniCPM-Vの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14、1024、64、96、2B、8B、2.5。
MiniCPM-Vの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連参照: https://arxiv.org/abs/2509.18154、https://huggingface.co/openbmb/MiniCPM-V-4.6、https://github.com/OpenBMB/MiniCPM-V。 値: 4.5、2025、4.6、2026、1B、400M、5-0.8B。
MouSi: Poly-Visual-Expert Vision-Language Models
MouSiの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
MouSiの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. MouSi integrates heterogeneous visual experts through a fusion network and projector into a language model. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
MouSiの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、558K。
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
InternLM-XComposer2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。
InternLM-XComposer2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. InternLM-XComposer2 applies Partial-LoRA only to visual tokens while preserving language-only behavior. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
InternLM-XComposer2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
MoE-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
MoE-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. MoE-LLaVA routes multimodal tokens through sparse experts added to a vision-encoder, projector, and language-model backbone. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
MoE-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
moondream1 and moondream2
moondream1 and moondream2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5。
Architecture figure: Official repository and model cards describe the implementation but provide no legitimate family architecture or system figure.
ℹ️ 詳細情報
moondream1 and moondream2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.6B、1.5、1.86B。
FireLLaVA
FireLLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 34B。
Architecture figure: A generic LLaVA diagram would not document FireLLaVA’s own contribution and should not be substituted.
ℹ️ 詳細情報
FireLLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 34B、588K。
COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
COSMOの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
COSMOの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 3. CosMo handles image and video inputs through a language model trained with contrastive and language-modeling objectives. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
COSMOの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 128、14。
TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
TinyGPT-Vの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。
TinyGPT-Vの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 4. TinyGPT-V projects frozen visual-backbone and Q-Former outputs through two linear layers into Phi-2. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
TinyGPT-Vの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、2.7。
MobileVLM: A Fast, Strong and Open Vision Language Assistant for Mobile Devices
MobileVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14。
MobileVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. MobileVLM connects a visual encoder to MobileLLaMA through a lightweight downsample projector. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
MobileVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14。
Alpha-CLIP: A CLIP Model Focusing on Wherever You Want
Alpha-CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Alpha-CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 3. Alpha-CLIP extends CLIP with an alpha channel that focuses visual encoding on a specified region. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
Alpha-CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 Alpha-CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 400M、5B。
Nous-Hermes-2-Vision - Mistral 7B
Nous-Hermes-2-Vision - Mistral 7Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2-、2.5、400M。
Nous-Hermes-2-Vision - Mistral 7Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Architecture figure: The card describes SigLIP integration and training data in prose but contains no legitimate architecture or system figure.
ℹ️ 詳細情報
Nous-Hermes-2-Vision - Mistral 7Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2-、2.5-、7B、400M、3B、220K、60K、150K、50K、2.5。
SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
SPHINXの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。
SPHINXの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 3. SPHINX jointly mixes tuning tasks, visual embeddings, and model weights in one multimodal architecture. Source paper, PDF p. 6. Figure notice.
ℹ️ 詳細情報
SPHINXの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、400M。
Florence-2: A Deep Dive into its Unified Architecture and Multi-Task Capabilities
Florence-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。
Bin Xiao, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng, Ce Liu, Lu Yuan
Figure 2. Florence-2 combines an image encoder and multimodality encoder-decoder through a unified task-prompt interface. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
Florence-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。
u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model
u-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
u-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. u-LLaVA unifies modality alignment with task-specific projectors, decoders, and patched downstream modules. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
u-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14、58K、23K。
LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
LLaVA-Plusの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
LLaVA-Plusの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. LLaVA-Plus retrieves skills, invokes tools, consumes their results, and generates a final response. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
LLaVA-Plusの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
OtterHD: A High-Resolution Multi-modality Model
OtterHDの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B。
OtterHDの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Architecture figure: The paper’s figures cover demonstrations, benchmark construction, throughput, resolution, and loss; none is a legitimate architecture substitute.
ℹ️ 詳細情報
OtterHDの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、2。
CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding
CoVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
CoVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. CoVLM vision module and communication-token framework. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
CoVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 97。
GLaMM: Pixel Grounding Large Multimodal Model
GLaMMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Hanoona Rasheed, Muhammad Maaz, Sahal Shaji Mullappilly, Abdelrahman Shaker, Salman Khan, Hisham Cholakkal, Rao M. Anwer, Erix Xing, Ming-Hsuan Yang, Fahad S. Khan
Figure 2. GLaMM architecture for scene-, region-, and pixel-level grounding. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
GLaMMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7.5、810。
Fuyu-8B: A Multimodal Architecture for AI Agents
Fuyu-8Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B。
Fuyu-8Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Official architecture diagram. Fuyu-8B projects image patches directly into a decoder-only Transformer. Primary source. Figure notice.
ℹ️ 詳細情報
Fuyu-8Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B。
PaLI-3 Vision Language Models: Smaller, Faster, Stronger
PaLI-3 Vision Language Modelsの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3、2B、3B。
PaLI-3 Vision Language Modelsの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. PaLI-3 connects a contrastively pretrained SigLIP encoder to an encoder-decoder UL2 Transformer. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
PaLI-3 Vision Language Modelsの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3、2B、3B。
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
MiniGPT-v2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、2-。
MiniGPT-v2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. MiniGPT-v2 architecture with frozen ViT, token concatenation, projection, and LLaMA-2. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
MiniGPT-v2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2-、7-、20M。
BakLLaVA
BakLLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、1.5、2、13B。
Architecture figure: The official model card and repository describe a LLaVA-on-Mistral derivative but publish no model-specific architecture figure.
ℹ️ 詳細情報
BakLLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、1.5、2、13B、600K、150K、558K、158K、450K、40K。
Ferret: Refer and Ground Anything Anywhere at Any Granularity
Ferretの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Ferretの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 3. Ferret hybrid region representation, spatial-aware sampler, and complete model architecture. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
Ferretの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.1。
LLaVA 1.5: Improved Baselines with Visual Instruction Tuning
LLaVA 1.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5。
LLaVA 1.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. LLaVA-1.5-HD grid-based encoding for arbitrary image resolutions. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
LLaVA 1.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5、11。
CogVLM: Visual Expert for Pretrained Language Models
CogVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
CogVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 4. CogVLM input pathway and visual-expert Transformer block. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
CogVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5、40、2B、700M。
MetaCLIP: Demystifying CLIP Data
MetaCLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
MetaCLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Architecture figure: The paper contributes data curation rather than a model architecture; Figure 5 is a data-pipeline case study.
ℹ️ 詳細情報
MetaCLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 400。
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Qwen-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Qwen-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 3. Qwen-VL visual-language architecture across pretraining, multitask pretraining, and supervised finetuning. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
Qwen-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
IDEFICS
IDEFICSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 80B、4。
Architecture figure: The official launch post and model card have capability and performance illustrations but no IDEFICS-specific architecture diagram.
ℹ️ 詳細情報
IDEFICSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 80、4。
BLIVA: A Simple Multimodal LLM for Better Handling of Text-rich Visual Questions
BLIVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
BLIVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. BLIVA architecture with frozen image encoder, Q-Former, patch projection, and frozen LLM. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
BLIVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
KOSMOS-2: Grounding Multimodal Large Language Models to the World
KOSMOS-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、1。
KOSMOS-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. KOSMOS-2 system overview for multimodal grounding and referring. Source paper, PDF p. 1. Figure notice.
ℹ️ 詳細情報
KOSMOS-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、1、256。
LaVIN: Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models
LaVINの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
LaVINの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. LaVIN architecture and Mixture-of-Modality Adaptation mechanism. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
LaVINの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
InstructBLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。
InstructBLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 3. InstructBLIP architecture with instruction-aware Q-Former and frozen LLM. Source paper, PDF p. 5. Figure notice.
ℹ️ 詳細情報
InstructBLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、26、11。
ImageBind: One Embedding Space To Bind Them All
ImageBindの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
ImageBindの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. ImageBind aligns six modalities in one shared embedding space. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
ImageBindの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
LLaVA: Large Language and Vision Assistant - Visual Instruction Tuning
LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. LLaVA connects CLIP visual features to a language model through a learned projection. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4、158K。
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
MiniGPT-4の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4。
MiniGPT-4の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. MiniGPT-4 architecture with ViT, Q-Former, linear projection, and Vicuna. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
MiniGPT-4の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4、20,000、256、3,500。
SigLIP: Sigmoid Loss for Language Image Pre-Training
SigLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas Beyer
Architecture figure: The paper changes the training loss, not the encoder architecture; Figure 1 is a distributed loss-implementation mock-up.
ℹ️ 詳細情報
SigLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
OpenFlamingoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14、7B。
OpenFlamingoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. OpenFlamingo-9B interleaved image-and-text system interface. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
OpenFlamingoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14、7-、7B、2B、64。
PaLM-E: An Embodied Multimodal Language Model
PaLM-Eの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
PaLM-Eの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. PaLM-E combines sensor encoders and a language model for embodied and visual-language tasks. Source paper, PDF p. 1. Figure notice.
ℹ️ 詳細情報
PaLM-Eの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
KOSMOS-1: Language Is Not All You Need: Aligning Perception with Language Models
KOSMOS-1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1。
KOSMOS-1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. KOSMOS-1 multimodal input, embedding, language-model, and output overview. Source paper, PDF p. 1. Figure notice.
ℹ️ 詳細情報
KOSMOS-1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1、2B、400M、700M。
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
BLIP-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。
BLIP-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. BLIP-2 bridges a frozen image encoder and frozen LLM through a two-stage Q-Former. Source paper, PDF p. 1. Figure notice.
ℹ️ 詳細情報
BLIP-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。
MULTIINSTRUCT: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
MULTIINSTRUCTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
MULTIINSTRUCTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Architecture figure: Figures cover examples, task taxonomy, performance, and attention; the paper publishes no model-specific architecture figure.
ℹ️ 詳細情報
MULTIINSTRUCTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
PaLI: A Jointly-Scaled Multilingual Language-Image Model
PaLIの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
PaLIの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. PaLI combines a scalable ViT with an encoder-decoder Transformer. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
PaLIの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4、10、100、17B。
Flamingo: a Visual Language Model for Few-Shot Learning
Flamingoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Flamingoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 3. Flamingo architecture for interleaved visual inputs and free-form text output. Source paper, PDF p. 4. Figure notice.
ℹ️ 詳細情報
Flamingoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
BLIP: Bootstrapping Language-Image Pre-training
BLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Junnan Li, Dongxu Li, Caiming Xiong, Steven Hoi
Figure 2. BLIP multimodal mixture-of-encoder-decoder architecture and training objectives. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
BLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 12M。
GLIP: Grounded Language-Image Pre-training
GLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
GLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. GLIP image and language encoders with deep fusion and word-region alignment. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
GLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
FROZEN: Multimodal Few-Shot Learning with Frozen Language Models
FROZENの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 50。
FROZENの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 2. FROZEN trains a vision encoder through a frozen language model. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
FROZENの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 50。
CLIP: Contrastive Language-Image Pre-training
CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 400。
CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. CLIP dual-encoder contrastive training and zero-shot classification approach. Source paper, PDF p. 2. Figure notice.
ℹ️ 詳細情報
CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 400。
ViT: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
ViTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
ViTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。
Figure 1. ViT patch embedding, Transformer encoder, and classification-token architecture. Source paper, PDF p. 3. Figure notice.
ℹ️ 詳細情報
ViTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 300M、100。