このページの内容

Awesome VLM Architectures

VLM Architecturesを扱う資料や関連プロジェクトをまとめたAwesomeリストです。

目次

引用

このリポジトリが役立つ場合は、以下の情報で引用できます。個別モデルの主張には原論文も引用してください。このカタログは文献案内であり、原資料の代替ではありません。

Gökay Aydoğanfal.aiで作成・保守しています(ORCIDgokay@fal.ai)。

構成図は図版クレジットで個別に帰属を示し、図版通知記載の権利に従います。

📚 BibTeX
@misc{aydogan2024awesomevlmarchitectures,
  author       = {Gökay Aydoğan},
  title        = {Awesome VLM Architectures},
  year         = {2024},
  howpublished = {\url{https://github.com/gokayfem/awesome-vlm-architectures}},
  note         = {GitHub repository, fal.ai},
  url          = {https://github.com/gokayfem/awesome-vlm-architectures}
}

モデル

すべてのアーキテクチャパネルはリリース日の新しい順です。同日公開のモデルは編集上のカタログ順を維持します。

🧭 Chronological Model Index (155 architectures, newest first)

2026: MODUS | Argus-Unified | Kimi K3 | Mage-VL | Inkling | Hy-Embodied-VLM | MonkeyOCRv2 | MiniMax M3 | InternVideo3 | Keye-VL 2.0 | Zamba2-VL | Cosmos 3 | Lance | ZAYA1-VL | Falcon Perception | GLM-5V-Turbo | PLaMo 2.1-VL | EXAONE 4.5 | BidirLM and BidirLM-Omni | Gemma 4 | Penguin-VL | Phi-4-Reasoning-Vision | V-SONAR and V-LCM | Qwen3.5 | Youtu-VL | Kimi K2.5 and K2.6 | Step3-VL-10B

2025: ERNIE 5.0 | DeepSeek-OCR | PaddleOCR-VL | Qwen3-VL | Step3 | GLM-4.1V-Thinking | ERNIE 4.5-VL | MiMo-VL | BAGEL | Seed1.5-VL | InternVL3 and InternVL3.5 | Kimi-VL | Llama 4 Scout and Maverick | Qwen2.5-Omni | Gemma 3 | Aya Vision | Phi-4-multimodal | SigLIP 2 | EVEv2 | Qwen2.5-VL | VideoLLaMA 3 | UI-TARS | MiniMax-01 | MiniCPM-o-2.6 | Eagle 2 | Sa2VA

2024: VideoChat-Flash | OmniVLM | Apollo | DeepSeek-VL2 | Maya | InternVL 2.5 | PaliGemma 2 | ShowUI | SmolVLM | AIMv2 | LLaVA-CoT | LLM2CLIP | Tarsier2 | Janus and Janus-Pro | ARIA | Emu3 | Molmo and PixMo | Llama 3.2-Vision | NVLM | Pixtral 12B | VILA-U | Qwen2-VL | EAGLE | Show-o | Idefics3-8B | Transfusion | mPLUG-Owl3 | VITA | LLaVA-OneVision | VILA² | INF-LLaVA | SlowFast-LLaVA | EVLM | InternLM-XComposer-2.5 | OMG-LLaVA | Cambrian-1 | EVE | Ovis | Parrot | ConvLLaVA | Phi-3-Vision and Phi-3.5-Vision | CogVLM2 | Chameleon | PaliGemma | xGen-MM (BLIP-3) | MANTIS | Moondream-next | Idefics2 | InternLM-XComposer2-4KHD | MM1 | DeepSeek-VL | AnyGPT | SPHINX-X | LLaVA 1.6 | MiniCPM-V | MouSi | InternLM-XComposer2 | MoE-LLaVA | moondream1 and moondream2 | FireLLaVA | COSMO

2023: TinyGPT-V | MobileVLM | Alpha-CLIP | Nous-Hermes-2-Vision - Mistral 7B | SPHINX | Florence-2 | u-LLaVA | LLaVA-Plus | OtterHD | CoVLM | GLaMM | Fuyu-8B | PaLI-3 Vision Language Models | MiniGPT-v2 | BakLLaVA | Ferret | LLaVA 1.5 | CogVLM | MetaCLIP | Qwen-VL | IDEFICS | BLIVA | KOSMOS-2 | LaVIN | InstructBLIP | ImageBind | LLaVA | MiniGPT-4 | SigLIP | OpenFlamingo | PaLM-E | KOSMOS-1 | BLIP-2

2022: MULTIINSTRUCT | PaLI | Flamingo | BLIP

2021: GLIP | FROZEN | CLIP

2020: ViT

リリース年表

日付は確認できる最初の公式モデルリリースを使い、ない場合は論文のarXiv v1投稿または最初の技術報告を使います。同系列のポイントリリースは最初のアーキテクチャ公開へ統合し、同日項目はカタログ順を維持します。

🗓️ Release Timeline (155 architectures, newest first)
日付アーキテクチャ特徴的な貢献
2026-07-28MODUSMODUSの特徴的なアーキテクチャ上の貢献
2026-07-28Argus-UnifiedArgus-Unifiedの特徴的なアーキテクチャ上の貢献
2026-07-27Kimi K3Kimi K3の特徴的なアーキテクチャ上の貢献
2026-07-27Mage-VLMage-VLの特徴的なアーキテクチャ上の貢献
2026-07-15InklingInklingの特徴的なアーキテクチャ上の貢献
2026-07-15Hy-Embodied-VLMHy-Embodied-VLMの特徴的なアーキテクチャ上の貢献
2026-07-11MonkeyOCRv2MonkeyOCRv2の特徴的なアーキテクチャ上の貢献
2026-06-11MiniMax M3MiniMax M3の特徴的なアーキテクチャ上の貢献
2026-06-10InternVideo3InternVideo3の特徴的なアーキテクチャ上の貢献
2026-06-09Keye-VL 2.0Keye-VL 2.0の特徴的なアーキテクチャ上の貢献
2026-06-02Zamba2-VLZamba2-VLの特徴的なアーキテクチャ上の貢献
2026-05-31Cosmos 3Cosmos 3の特徴的なアーキテクチャ上の貢献
2026-05-18LanceLanceの特徴的なアーキテクチャ上の貢献
2026-05-08ZAYA1-VLZAYA1-VLの特徴的なアーキテクチャ上の貢献
2026-05-03Falcon PerceptionFalcon Perceptionの特徴的なアーキテクチャ上の貢献
2026-04-29GLM-5V-TurboGLM-5V-Turboの特徴的なアーキテクチャ上の貢献
2026-04-21PLaMo 2.1-VLPLaMo 2.1-VLの特徴的なアーキテクチャ上の貢献
2026-04-09EXAONE 4.5EXAONE 4.5の特徴的なアーキテクチャ上の貢献
2026-04-02BidirLM and BidirLM-OmniBidirLM and BidirLM-Omniの特徴的なアーキテクチャ上の貢献
2026-03-31Gemma 4Gemma 4の特徴的なアーキテクチャ上の貢献
2026-03-06Penguin-VLPenguin-VLの特徴的なアーキテクチャ上の貢献
2026-03-04Phi-4-Reasoning-VisionPhi-4-Reasoning-Visionの特徴的なアーキテクチャ上の貢献
2026-03-01V-SONAR and V-LCMV-SONAR and V-LCMの特徴的なアーキテクチャ上の貢献
2026-02-16Qwen3.5Qwen3.5の特徴的なアーキテクチャ上の貢献
2026-01-27Youtu-VLYoutu-VLの特徴的なアーキテクチャ上の貢献
2026-01-27Kimi K2.5 and K2.6Kimi K2.5 and K2.6の特徴的なアーキテクチャ上の貢献
2026-01-14Step3-VL-10BStep3-VL-10Bの特徴的なアーキテクチャ上の貢献
2025-11-13ERNIE 5.0ERNIE 5.0の特徴的なアーキテクチャ上の貢献
2025-10-20DeepSeek-OCRDeepSeek-OCRの特徴的なアーキテクチャ上の貢献
2025-10-16PaddleOCR-VLPaddleOCR-VLの特徴的なアーキテクチャ上の貢献
2025-09-22Qwen3-VLQwen3-VLの特徴的なアーキテクチャ上の貢献
2025-07-25Step3Step3の特徴的なアーキテクチャ上の貢献
2025-07-01GLM-4.1V-ThinkingGLM-4.1V-Thinkingの特徴的なアーキテクチャ上の貢献
2025-06-30ERNIE 4.5-VLERNIE 4.5-VLの特徴的なアーキテクチャ上の貢献
2025-06-04MiMo-VLMiMo-VLの特徴的なアーキテクチャ上の貢献
2025-05-20BAGELBAGELの特徴的なアーキテクチャ上の貢献
2025-05-11Seed1.5-VLSeed1.5-VLの特徴的なアーキテクチャ上の貢献
2025-04-11InternVL3 and InternVL3.5InternVL3 and InternVL3.5の特徴的なアーキテクチャ上の貢献
2025-04-10Kimi-VLKimi-VLの特徴的なアーキテクチャ上の貢献
2025-04-05Llama 4 Scout and MaverickLlama 4 Scout and Maverickの特徴的なアーキテクチャ上の貢献
2025-03-26Qwen2.5-OmniQwen2.5-Omniの特徴的なアーキテクチャ上の貢献
2025-03-12Gemma 3Gemma 3の特徴的なアーキテクチャ上の貢献
2025-03-04Aya VisionAya Visionの特徴的なアーキテクチャ上の貢献
2025-03-03Phi-4-multimodalPhi-4-multimodalの特徴的なアーキテクチャ上の貢献
2025-02-20SigLIP 2SigLIP 2の特徴的なアーキテクチャ上の貢献
2025-02-08EVEv2EVEv2の特徴的なアーキテクチャ上の貢献
2025-01-26Qwen2.5-VLQwen2.5-VLの特徴的なアーキテクチャ上の貢献
2025-01-21VideoLLaMA 3VideoLLaMA 3の特徴的なアーキテクチャ上の貢献
2025-01-20UI-TARSUI-TARSの特徴的なアーキテクチャ上の貢献
2025-01-14MiniMax-01MiniMax-01の特徴的なアーキテクチャ上の貢献
2025-01-12MiniCPM-o-2.6MiniCPM-o-2.6の特徴的なアーキテクチャ上の貢献
2025-01-10Eagle 2Eagle 2の特徴的なアーキテクチャ上の貢献
2025-01-07Sa2VASa2VAの特徴的なアーキテクチャ上の貢献
2024-12-31VideoChat-FlashVideoChat-Flashの特徴的なアーキテクチャ上の貢献
2024-12-16OmniVLMOmniVLMの特徴的なアーキテクチャ上の貢献
2024-12-13ApolloApolloの特徴的なアーキテクチャ上の貢献
2024-12-13DeepSeek-VL2DeepSeek-VL2の特徴的なアーキテクチャ上の貢献
2024-12-10MayaMayaの特徴的なアーキテクチャ上の貢献
2024-12-05InternVL 2.5InternVL 2.5の特徴的なアーキテクチャ上の貢献
2024-12-04PaliGemma 2PaliGemma 2の特徴的なアーキテクチャ上の貢献
2024-11-26ShowUIShowUIの特徴的なアーキテクチャ上の貢献
2024-11-26SmolVLMSmolVLMの特徴的なアーキテクチャ上の貢献
2024-11-21AIMv2AIMv2の特徴的なアーキテクチャ上の貢献
2024-11-15LLaVA-CoTLLaVA-CoTの特徴的なアーキテクチャ上の貢献
2024-11-06LLM2CLIPLLM2CLIPの特徴的なアーキテクチャ上の貢献
2024-11-05Tarsier2Tarsier2の特徴的なアーキテクチャ上の貢献
2024-10-17Janus and Janus-ProJanus and Janus-Proの特徴的なアーキテクチャ上の貢献
2024-10-08ARIAARIAの特徴的なアーキテクチャ上の貢献
2024-09-27Emu3Emu3の特徴的なアーキテクチャ上の貢献
2024-09-25Molmo and PixMoMolmo and PixMoの特徴的なアーキテクチャ上の貢献
2024-09-25Llama 3.2-VisionLlama 3.2-Visionの特徴的なアーキテクチャ上の貢献
2024-09-17NVLMNVLMの特徴的なアーキテクチャ上の貢献
2024-09-11Pixtral 12BPixtral 12Bの特徴的なアーキテクチャ上の貢献
2024-09-06VILA-UVILA-Uの特徴的なアーキテクチャ上の貢献
2024-08-29Qwen2-VLQwen2-VLの特徴的なアーキテクチャ上の貢献
2024-08-28EAGLEEAGLEの特徴的なアーキテクチャ上の貢献
2024-08-22Show-oShow-oの特徴的なアーキテクチャ上の貢献
2024-08-22Idefics3-8BIdefics3-8Bの特徴的なアーキテクチャ上の貢献
2024-08-20TransfusionTransfusionの特徴的なアーキテクチャ上の貢献
2024-08-09mPLUG-Owl3mPLUG-Owl3の特徴的なアーキテクチャ上の貢献
2024-08-09VITAVITAの特徴的なアーキテクチャ上の貢献
2024-08-05LLaVA-OneVisionLLaVA-OneVisionの特徴的なアーキテクチャ上の貢献
2024-07-24VILA²VILA²の特徴的なアーキテクチャ上の貢献
2024-07-23INF-LLaVAINF-LLaVAの特徴的なアーキテクチャ上の貢献
2024-07-22SlowFast-LLaVASlowFast-LLaVAの特徴的なアーキテクチャ上の貢献
2024-07-19EVLMEVLMの特徴的なアーキテクチャ上の貢献
2024-07-03InternLM-XComposer-2.5InternLM-XComposer-2.5の特徴的なアーキテクチャ上の貢献
2024-06-27OMG-LLaVAOMG-LLaVAの特徴的なアーキテクチャ上の貢献
2024-06-24Cambrian-1Cambrian-1の特徴的なアーキテクチャ上の貢献
2024-06-17EVEEVEの特徴的なアーキテクチャ上の貢献
2024-06-14OvisOvisの特徴的なアーキテクチャ上の貢献
2024-06-04ParrotParrotの特徴的なアーキテクチャ上の貢献
2024-05-24ConvLLaVAConvLLaVAの特徴的なアーキテクチャ上の貢献
2024-05-21Phi-3-Vision and Phi-3.5-VisionPhi-3-Vision and Phi-3.5-Visionの特徴的なアーキテクチャ上の貢献
2024-05-20CogVLM2CogVLM2の特徴的なアーキテクチャ上の貢献
2024-05-16ChameleonChameleonの特徴的なアーキテクチャ上の貢献
2024-05-14PaliGemmaPaliGemmaの特徴的なアーキテクチャ上の貢献
2024-05-06xGen-MM (BLIP-3)xGen-MM (BLIP-3)の特徴的なアーキテクチャ上の貢献
2024-05-02MANTISMANTISの特徴的なアーキテクチャ上の貢献
2024-04-19Moondream-nextMoondream-nextの特徴的なアーキテクチャ上の貢献
2024-04-15Idefics2Idefics2の特徴的なアーキテクチャ上の貢献
2024-04-09InternLM-XComposer2-4KHDInternLM-XComposer2-4KHDの特徴的なアーキテクチャ上の貢献
2024-03-14MM1MM1の特徴的なアーキテクチャ上の貢献
2024-03-08DeepSeek-VLDeepSeek-VLの特徴的なアーキテクチャ上の貢献
2024-02-19AnyGPTAnyGPTの特徴的なアーキテクチャ上の貢献
2024-02-08SPHINX-XSPHINX-Xの特徴的なアーキテクチャ上の貢献
2024-01-30LLaVA 1.6LLaVA 1.6の特徴的なアーキテクチャ上の貢献
2024-01-30MiniCPM-VMiniCPM-Vの特徴的なアーキテクチャ上の貢献
2024-01-30MouSiMouSiの特徴的なアーキテクチャ上の貢献
2024-01-29InternLM-XComposer2InternLM-XComposer2の特徴的なアーキテクチャ上の貢献
2024-01-29MoE-LLaVAMoE-LLaVAの特徴的なアーキテクチャ上の貢献
2024-01-20moondream1 and moondream2moondream1 and moondream2の特徴的なアーキテクチャ上の貢献
2024-01-05FireLLaVAFireLLaVAの特徴的なアーキテクチャ上の貢献
2024-01-01COSMOCOSMOの特徴的なアーキテクチャ上の貢献
2023-12-28TinyGPT-VTinyGPT-Vの特徴的なアーキテクチャ上の貢献
2023-12-28MobileVLMMobileVLMの特徴的なアーキテクチャ上の貢献
2023-12-06Alpha-CLIPAlpha-CLIPの特徴的なアーキテクチャ上の貢献
2023-11-28Nous-Hermes-2-Vision - Mistral 7BNous-Hermes-2-Vision - Mistral 7Bの特徴的なアーキテクチャ上の貢献
2023-11-13SPHINXSPHINXの特徴的なアーキテクチャ上の貢献
2023-11-10Florence-2Florence-2の特徴的なアーキテクチャ上の貢献
2023-11-09u-LLaVAu-LLaVAの特徴的なアーキテクチャ上の貢献
2023-11-09LLaVA-PlusLLaVA-Plusの特徴的なアーキテクチャ上の貢献
2023-11-07OtterHDOtterHDの特徴的なアーキテクチャ上の貢献
2023-11-06CoVLMCoVLMの特徴的なアーキテクチャ上の貢献
2023-11-06GLaMMGLaMMの特徴的なアーキテクチャ上の貢献
2023-10-17Fuyu-8BFuyu-8Bの特徴的なアーキテクチャ上の貢献
2023-10-13PaLI-3 Vision Language ModelsPaLI-3 Vision Language Modelsの特徴的なアーキテクチャ上の貢献
2023-10-13MiniGPT-v2MiniGPT-v2の特徴的なアーキテクチャ上の貢献
2023-10-12BakLLaVABakLLaVAの特徴的なアーキテクチャ上の貢献
2023-10-11FerretFerretの特徴的なアーキテクチャ上の貢献
2023-10-05LLaVA 1.5LLaVA 1.5の特徴的なアーキテクチャ上の貢献
2023-10-05CogVLMCogVLMの特徴的なアーキテクチャ上の貢献
2023-09-28MetaCLIPMetaCLIPの特徴的なアーキテクチャ上の貢献
2023-08-24Qwen-VLQwen-VLの特徴的なアーキテクチャ上の貢献
2023-08-22IDEFICSIDEFICSの特徴的なアーキテクチャ上の貢献
2023-08-19BLIVABLIVAの特徴的なアーキテクチャ上の貢献
2023-06-26KOSMOS-2KOSMOS-2の特徴的なアーキテクチャ上の貢献
2023-05-24LaVINLaVINの特徴的なアーキテクチャ上の貢献
2023-05-11InstructBLIPInstructBLIPの特徴的なアーキテクチャ上の貢献
2023-05-09ImageBindImageBindの特徴的なアーキテクチャ上の貢献
2023-04-17LLaVALLaVAの特徴的なアーキテクチャ上の貢献
2023-04-16MiniGPT-4MiniGPT-4の特徴的なアーキテクチャ上の貢献
2023-03-27SigLIPSigLIPの特徴的なアーキテクチャ上の貢献
2023-03-14OpenFlamingoOpenFlamingoの特徴的なアーキテクチャ上の貢献
2023-03-06PaLM-EPaLM-Eの特徴的なアーキテクチャ上の貢献
2023-02-27KOSMOS-1KOSMOS-1の特徴的なアーキテクチャ上の貢献
2023-01-30BLIP-2BLIP-2の特徴的なアーキテクチャ上の貢献
2022-12-21MULTIINSTRUCTMULTIINSTRUCTの特徴的なアーキテクチャ上の貢献
2022-09-14PaLIPaLIの特徴的なアーキテクチャ上の貢献
2022-04-28FlamingoFlamingoの特徴的なアーキテクチャ上の貢献
2022-01-28BLIPBLIPの特徴的なアーキテクチャ上の貢献
2021-12-07GLIPGLIPの特徴的なアーキテクチャ上の貢献
2021-06-25FROZENFROZENの特徴的なアーキテクチャ上の貢献
2021-01-05CLIPCLIPの特徴的なアーキテクチャ上の貢献
2020-10-22ViTViTの特徴的なアーキテクチャ上の貢献

アーキテクチャ

MODUS: Decoder-Only Any-to-Any Multimodal Modeling

MODUSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv Project

Mingqiao Ye et al., EPFL
Released: 2026-07-28

MODUS: Decoder-Only Any-to-Any Multimodal Modeling architecture: Decoder-only any-to-any modeling across tokenized 1D and 2D modalities

Figure 2. Decoder-only any-to-any modeling across tokenized 1D and 2D modalities. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

MODUSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

MODUSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Argus-Unified: Economical Understanding and Generation

Argus-Unifiedの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv

Weiming Zhuang et al.
Released: 2026-07-28

Argus-Unified: Economical Understanding and Generation architecture: Two-stage hybrid-token training for unified image understanding and generation

Figure 3. Two-stage hybrid-token training for unified image understanding and generation. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

Argus-Unifiedの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Argus-Unifiedの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 15.6、2,000。

Kimi K3: Kimi Delta Attention at Trillion-Parameter Scale

Kimi K3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 896。

arXiv GitHub HuggingFace

Kimi Team, Moonshot AI
Released: 2026-07-27

Kimi K3: Kimi Delta Attention at Trillion-Parameter Scale architecture: Kimi Delta Attention, Stable LatentMoE, Attention Residuals, and MoonViT-V2

Figure 2. Kimi Delta Attention, Stable LatentMoE, Attention Residuals, and MoonViT-V2. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

Kimi K3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.8T、104B、93、24、16、896。

Kimi K3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 401M、1,048,576-。

Mage-VL: Codec-Native Streaming Multimodality

Mage-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv Project

Senqiao Yang et al., Microsoft Research
Released: 2026-07-27

Mage-VL: Codec-Native Streaming Multimodality architecture: Codec-native streaming perception with an event gate and causal language decoder

Figure 3. Codec-native streaming perception with an event gate and causal language decoder. Source paper, PDF p. 8. Figure notice.

ℹ️ 詳細情報

Mage-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 16×16-、75、560M、100M。

Mage-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3.5。

Inkling: Relative-Position Multimodal Mixture of Experts

Inklingの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Website Model Card HuggingFace

Thinking Machines Lab
Released: 2026-07-15

Architecture figure: The official Inkling model card contains no architecture figure.

ℹ️ 詳細情報

Inklingの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 975B、41B、256。

Inklingの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 45T。

Hy-Embodied-VLM: Sparse-MoE Reasoning for Physical Agents

Hy-Embodied-VLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.0。

arXiv GitHub HuggingFace

Tencent Robotics X, Hy Vision Team and Futian Laboratory
Released: 2026-07-15

Hy-Embodied-VLM: Sparse-MoE Reasoning for Physical Agents architecture: Self-evolving supervised fine-tuning, rejection sampling, and specialized reinforcement learning

Figure 4. Self-evolving supervised fine-tuning, rejection sampling, and specialized reinforcement learning. Source paper, PDF p. 11. Figure notice.

ℹ️ 詳細情報

Hy-Embodied-VLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.0、30B、3B、128、32K。

Hy-Embodied-VLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

MonkeyOCRv2: Document-Native Visual-Text Pretraining

MonkeyOCRv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 17。

arXiv GitHub

Yuliang Liu et al.
Released: 2026-07-11

MonkeyOCRv2: Document-Native Visual-Text Pretraining architecture: Document-native pretraining through text generation and pixel reconstruction

Figure 1. Document-native pretraining through text generation and pixel reconstruction. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

MonkeyOCRv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

MonkeyOCRv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 113M、17、0.7B、11。

MiniMax M3: Native Multimodality with Sparse Long-Context Attention

MiniMax M3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 109B。

arXiv Website HuggingFace

MiniMax
Released: 2026-06-11

MiniMax M3: Native Multimodality with Sparse Long-Context Attention architecture: MiniMax Sparse Attention index and exact-attention branches

Figure 1. MiniMax Sparse Attention index and exact-attention branches. Source paper, PDF p. 1. Figure notice.

ℹ️ 詳細情報

MiniMax M3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

MiniMax M3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 109B。

InternVideo3: Multimodal Contextual Reasoning for Video Agents

InternVideo3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv

Ziang Yan et al.
Released: 2026-06-10

InternVideo3: Multimodal Contextual Reasoning for Video Agents architecture: InternVideo3 with multimodal multi-head latent attention across long contexts

Figure 2. InternVideo3 with multimodal multi-head latent attention across long contexts. Source paper, PDF p. 7. Figure notice.

ℹ️ 詳細情報

InternVideo3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

InternVideo3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Keye-VL 2.0: Sparse Attention for Long-Video Agents

Keye-VL 2.0の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.0、256K。

arXiv

Kwai Keye Team
Released: 2026-06-09

Keye-VL 2.0: Sparse Attention for Long-Video Agents architecture: Four-stage curriculum extending Keye-VL from alignment to 256K context

Figure 2. Four-stage curriculum extending Keye-VL from alignment to 256K context. Source paper, PDF p. 8. Figure notice.

ℹ️ 詳細情報

Keye-VL 2.0の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.0-30B、3B、30B、256K。

Keye-VL 2.0の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Zamba2-VL: Hybrid State-Space Vision-Language Modeling

Zamba2-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。

arXiv Project

Zyphra
Released: 2026-06-02

Zamba2-VL: Hybrid State-Space Vision-Language Modeling architecture: Zamba2 hybrid state-space language backbone connected to a vision encoder

Figure 1. Zamba2 hybrid state-space language backbone connected to a vision encoder. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

Zamba2-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.2B、2.7B、7B、2。

Zamba2-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Cosmos 3: Omnimodal World Modeling with Mixture of Transformers

Cosmos 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3。

arXiv GitHub HuggingFace

NVIDIA
Released: 2026-05-31

Cosmos 3: Omnimodal World Modeling with Mixture of Transformers architecture: Mixture-of-Transformers reasoner and generator with shared attention

Figure 5. Mixture-of-Transformers reasoner and generator with shared attention. Source paper, PDF p. 11. Figure notice.

ℹ️ 詳細情報

Cosmos 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3。

Cosmos 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、32B、20、3。

Lance: Unified Image and Video Understanding, Generation, and Editing

Lanceの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv Project

Lance Team
Released: 2026-05-18

Lance: Unified Image and Video Understanding, Generation, and Editing architecture: Dual-expert sequence modeling for understanding and visual generation

Figure 6. Dual-expert sequence modeling for understanding and visual generation. Source paper, PDF p. 9. Figure notice.

ℹ️ 詳細情報

Lanceの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3B。

Lanceの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 128。

ZAYA1-VL: Vision-Specialized Compressed Convolutional Attention

ZAYA1-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv Project HuggingFace

Zyphra
Released: 2026-05-08

ZAYA1-VL: Vision-Specialized Compressed Convolutional Attention architecture: Visual routing, compressed convolutional attention, and a hybrid language backbone

Figure 2. Visual routing, compressed convolutional attention, and a hybrid language backbone. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

ZAYA1-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、5-。

ZAYA1-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 140B、2.0。

Falcon Perception: Early-Fusion Grounding, Segmentation, and OCR

Falcon Perceptionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv Project Release

Technology Innovation Institute
Released: 2026-05-03

Falcon Perception: Early-Fusion Grounding, Segmentation, and OCR architecture: Early-fusion perception Transformer with grounding, geometry, and segmentation pathways

Figure 1. Early-fusion perception Transformer with grounding, geometry, and segmentation pathways. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

Falcon Perceptionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Falcon Perceptionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 600M、300M、28、3。

GLM-5V-Turbo: Native Multimodal Agency

GLM-5V-Turboの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv

GLM-V Team
Released: 2026-04-29

GLM-5V-Turbo: Native Multimodal Agency architecture: Multimodal multi-token prediction with image placeholders and shared Transformer blocks

Figure 2. Multimodal multi-token prediction with image placeholders and shared Transformer blocks. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

GLM-5V-Turboの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

GLM-5V-Turboの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.。

PLaMo 2.1-VL: Lightweight Japanese Vision-Language Modeling

PLaMo 2.1-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.1-、2B、8B。

arXiv

Tommi Kerola et al., Preferred Networks
Released: 2026-04-21

Architecture figure: The PLaMo 2.1-VL paper contains application and data figures, but no model architecture diagram.

ℹ️ 詳細情報

PLaMo 2.1-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.1-。

PLaMo 2.1-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

EXAONE 4.5: Native Multimodal Pretraining for Documents

EXAONE 4.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.5。

arXiv

Eunbi Choi et al., LG AI Research
Released: 2026-04-09

EXAONE 4.5: Native Multimodal Pretraining for Documents architecture: Native-resolution vision encoding, projection, language decoding, and multi-token prediction

Figure 1. Native-resolution vision encoding, projection, language decoding, and multi-token prediction. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

EXAONE 4.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.5、4.0。

EXAONE 4.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 256K。

BidirLM and BidirLM-Omni: Causal Decoders as Multimodal Encoders

BidirLM and BidirLM-Omniの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv HuggingFace Release

BidirLM Team
Released: 2026-04-02

BidirLM and BidirLM-Omni: Causal Decoders as Multimodal Encoders architecture: Specialist-backbone merging with frozen modality projection heads

Figure 10. Specialist-backbone merging with frozen modality projection heads. Source paper, PDF p. 25. Figure notice.

ℹ️ 詳細情報

BidirLM and BidirLM-Omniの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

BidirLM and BidirLM-Omniの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Gemma 4: Open-Weight Native Multimodal Models

Gemma 4の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4。

arXiv Releases HuggingFace

Gemma Team, Google DeepMind
Released: 2026-03-31

Gemma 4: Open-Weight Native Multimodal Models architecture: Aspect-preserving image resizing, patch pooling, and soft-token production

Figure 2. Aspect-preserving image resizing, patch pooling, and soft-token production. Source paper, PDF p. 16. Figure notice.

ℹ️ 詳細情報

Gemma 4の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4、31B、26B、4B、12B、128K、256K。

Gemma 4の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Gemma 4の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Penguin-VL: Efficient VLMs with LLM-Based Vision Encoders

Penguin-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace

Boqiang Zhang, Lei Ke, Ruihan Yang, Qi Gao, Tianyuan Qu, Rossell Chen, Dong Yu, Leoweiliang
Released: 2026-03-06

Penguin-VL: Efficient VLMs with LLM-Based Vision Encoders architecture: An LLM-initialized vision encoder with priority-aware video-token compression

Figure 3. An LLM-initialized vision encoder with priority-aware video-token compression. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

Penguin-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 0.6B、2B、8B。

Penguin-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Penguin-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Phi-4-Reasoning-Vision: Compact Multimodal Reasoning

Phi-4-Reasoning-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4-、15B。

arXiv GitHub HuggingFace

Jyoti Aneja, Michael Harrison, Neel Joshi, Tyler LaBonte, John Langford, Eduardo Salinas
Released: 2026-03-04

Phi-4-Reasoning-Vision: Compact Multimodal Reasoning architecture: SigLIP2 vision encoding, cross-modal projection, mid-fusion, and language reasoning

Figure 3. SigLIP2 vision encoding, cross-modal projection, mid-fusion, and language reasoning. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

Phi-4-Reasoning-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4-、15B、2、3,600。

Phi-4-Reasoning-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連コード: <think><nothink>。 値: 240。

Phi-4-Reasoning-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

V-SONAR and V-LCM: Vision-Language Modeling in Concept Space

V-SONAR and V-LCMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv OpenReview

Yifu Qiu, Paul-Ambroise Duquenne, Holger Schwenk
Released: 2026-03-01

V-SONAR and V-LCM: Vision-Language Modeling in Concept Space architecture: Visual-semantic alignment and concept-space prediction

Figure 1. Visual-semantic alignment and concept-space prediction. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

V-SONAR and V-LCMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

V-SONAR and V-LCMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

V-SONAR and V-LCMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 80、1K。

Qwen3.5: Native Multimodal Hybrid-Attention Models

Qwen3.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、6。

Blog GitHub HuggingFace

Qwen Team
Released: 2026-02-16

Architecture figure: Qwen3.5 has no public technical paper containing an architecture figure.

ℹ️ 詳細情報

Qwen3.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、397B、17B。

Qwen3.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Qwen3.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 6、5。

Youtu-VL: Unified Autoregressive Supervision for Dense Vision

Youtu-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace

Tencent Youtu Lab
Released: 2026-01-27

Youtu-VL: Unified Autoregressive Supervision for Dense Vision architecture: Unified visual-text autoregressive supervision and dense-output decoding

Figure 3. Unified visual-text autoregressive supervision and dense-output decoding. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

Youtu-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Youtu-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Kimi K2.5 and K2.6: Native Multimodal Agentic MoE

Kimi K2.5 and K2.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、6。

arXiv HuggingFace Release

Kimi Team, Moonshot AI
Released: 2026-01-27

Kimi K2.5 and K2.6: Native Multimodal Agentic MoE architecture: Agentic reinforcement-learning environments, rollout management, and training services

Figure 10. Agentic reinforcement-learning environments, rollout management, and training services. Source paper, PDF p. 23. Figure notice.

ℹ️ 詳細情報

Kimi K2.5 and K2.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、1T、32B、61、384。

Kimi K2.5 and K2.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 20、6、5、256K。

Step3-VL-10B: Language-Aligned Perception with 16× Token Compression

Step3-VL-10Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 10B、1.8B、8B。

arXiv Project

StepFun
Released: 2026-01-14

Architecture figure: The Step3-VL-10B report contains performance and RL figures, but no architecture diagram.

ℹ️ 詳細情報

Step3-VL-10Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 10B、2025、1.8B、2、16、8B。

Step3-VL-10Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 728×728、504×504。

ERNIE 5.0: Unified Autoregressive Omnimodal Mixture of Experts

ERNIE 5.0の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5.0、2.4T。

arXiv Architecture

Baidu ERNIE Team
Released: 2025-11-13

ERNIE 5.0: Unified Autoregressive Omnimodal Mixture of Experts architecture: Unified image understanding, image generation, and video generation objectives

Figure 2. Unified image understanding, image generation, and video generation objectives. Source paper, PDF p. 6. Figure notice.

ℹ️ 詳細情報

ERNIE 5.0の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5.0、3、2.4T。

ERNIE 5.0の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5.0、13、2025、2026、4。

DeepSeek-OCR: Visual Context Compression through DeepEncoder

DeepSeek-OCRの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace

DeepSeek-AI
Released: 2025-10-20

DeepSeek-OCR: Visual Context Compression through DeepEncoder architecture: A SAM-CLIP DeepEncoder connected to a sparse language decoder

Figure 3. A SAM-CLIP DeepEncoder connected to a sparse language decoder. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

DeepSeek-OCRの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3B、570M、64、800。

DeepSeek-OCRの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

DeepSeek-OCRの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連参照: https://arxiv.org/abs/2601.20552https://github.com/deepseek-ai/DeepSeek-OCR-2。 値: 2、27、2026。

PaddleOCR-VL: Ultra-Compact Multilingual Document Parsing

PaddleOCR-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.5-0.3B、0.9B。

arXiv GitHub HuggingFace

PaddleOCR Team, Baidu
Released: 2025-10-16

PaddleOCR-VL: Ultra-Compact Multilingual Document Parsing architecture: Document layout analysis, compact VLM inference, instructions, and structured output

Figure 2. Document layout analysis, compact VLM inference, instructions, and structured output. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

PaddleOCR-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 0.9B、4.5-0.3B、109。

PaddleOCR-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5、2026、1.6。

Qwen3-VL: DeepStack Vision-Language Models

Qwen3-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace

Qwen Team
Released: 2025-09-22

Qwen3-VL: DeepStack Vision-Language Models architecture: Vision encoding, DeepStack injection, and dense or mixture-of-experts decoding

Figure 1. Vision encoding, DeepStack injection, and dense or mixture-of-experts decoding. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

Qwen3-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2B、4B、8B、32B、30B、235B、256K。

Qwen3-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Qwen3-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Step3: Model-System Co-Design for Cost-Effective Multimodal Intelligence

Step3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 321B、38B。

arXiv GitHub Website

StepFun
Released: 2025-07-25

Step3: Model-System Co-Design for Cost-Effective Multimodal Intelligence architecture: Attention-FFN disaggregation across attention and expert instances

Figure 6. Attention-FFN disaggregation across attention and expert instances. Source paper, PDF p. 11. Figure notice.

ℹ️ 詳細情報

Step3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 321B、38B。

Step3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

GLM-4.1V-Thinking: General-Purpose Multimodal Reasoning through Curriculum-Sampled RL

GLM-4.1V-Thinkingの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.、9B。

arXiv GitHub HuggingFace

GLM-V Team, Zhipu AI and Tsinghua University
Released: 2025-07-01

GLM-4.1V-Thinking: General-Purpose Multimodal Reasoning through Curriculum-Sampled RL architecture: Native-resolution vision encoding, projection, decoding, and timestamped video tokens

Figure 2. Native-resolution vision encoding, projection, decoding, and timestamped video tokens. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

GLM-4.1V-Thinkingの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.、9B、4-9B、0414。

GLM-4.1V-Thinkingの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連コード: <think><answer>。 値: 32K。

GLM-4.1V-Thinkingの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 220M、4.。

ERNIE 4.5-VL: Heterogeneous Modality Mixture-of-Experts

ERNIE 4.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.5-。

Website GitHub HuggingFace

Baidu ERNIE Team
Released: 2025-06-30

Architecture figure: ERNIE 4.5-VL has no public paper containing an extractable architecture figure.

ℹ️ 詳細情報

ERNIE 4.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.5-、424B、47B、28B、3B、4.5。

ERNIE 4.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

MiMo-VL: Multimodal Pretraining with Mixed On-Policy Reinforcement Learning

MiMo-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.4T。

arXiv GitHub HuggingFace

Xiaomi MiMo Team
Released: 2025-06-04

MiMo-VL: Multimodal Pretraining with Mixed On-Policy Reinforcement Learning architecture: Native-resolution vision encoding, projection, and language decoding

Figure 2. Native-resolution vision encoding, projection, and language decoding. Source paper, PDF p. 6. Figure notice.

ℹ️ 詳細情報

MiMo-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B。

MiMo-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.4T。

BAGEL: A Mixture-of-Transformer-Experts for Unified Understanding and Generation

BAGELの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace

Chaorui Deng, Deyao Zhu, Kunchang Li, Chenhui Gou, Feng Li, Zeyu Wang, Shu Zhong, Weihao Yu, Xiaonan Nie, Ziang Song, Guang Shi, Haoqi Fan
Released: 2025-05-20

BAGEL: A Mixture-of-Transformer-Experts for Unified Understanding and Generation architecture: Shared self-attention with understanding and generation Transformer experts

Figure 2. Shared self-attention with understanding and generation Transformer experts. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

BAGELの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14B、7B、5-、2。

BAGELの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

BAGELの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 400M、500M、1.6B、100M、45M、20M、500K。

Seed1.5-VL: Sparse-MoE Multimodal Understanding and Agentic Reasoning

Seed1.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5-、532M、20B。

arXiv GitHub

ByteDance Seed Team
Released: 2025-05-11

Seed1.5-VL: Sparse-MoE Multimodal Understanding and Agentic Reasoning architecture: Native-resolution vision encoding, adaptation, sparse MoE decoding, and timestamped video

Figure 1. Native-resolution vision encoding, adaptation, sparse MoE decoding, and timestamped video. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

Seed1.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5-、20B。

Seed1.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

InternVL3 and InternVL3.5: Native Multimodal Pretraining and Adaptive Resolution

InternVL3 and InternVL3.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5。

arXiv arXiv GitHub HuggingFace

InternVL Team, OpenGVLab
Released: 2025-04-11

Architecture figure: The InternVL3 paper contains evaluation figures, but no definitive model architecture diagram.

ℹ️ 詳細情報

InternVL3 and InternVL3.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

InternVL3 and InternVL3.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 26、2025、5。

Kimi-VL: Native-Resolution Vision with a Sparse MoE Decoder

Kimi-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.8B、16B、128K。

arXiv GitHub HuggingFace

Kimi Team
Released: 2025-04-10

Kimi-VL: Native-Resolution Vision with a Sparse MoE Decoder architecture: MoonViT, multimodal projection, and a sparse mixture-of-experts decoder

Figure 3. MoonViT, multimodal projection, and a sparse mixture-of-experts decoder. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

Kimi-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 16B、2.8B、2×2。

Kimi-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5.2T、4.4T、2T、0.1T、2.3T、8K、128K、2506。

Kimi-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Llama 4 Scout and Maverick: Native Multimodal Mixture-of-Experts Models

Llama 4 Scout and Maverickの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4、17B。

Model Card HuggingFace

Meta
Released: 2025-04-05

Architecture figure: The official Llama 4 model card describes the architecture in text and tables only.

ℹ️ 詳細情報

Llama 4 Scout and Maverickの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4、109B、16、400B、128、17B、10M、1M。

Llama 4 Scout and Maverickの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 40T、22T、4。

Qwen2.5-Omni: Streaming Multimodal Perception and Speech Generation

Qwen2.5-Omniの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5-。

arXiv GitHub HuggingFace

Qwen Team
Released: 2025-03-26

Qwen2.5-Omni: Streaming Multimodal Perception and Speech Generation architecture: Thinker-Talker architecture for multimodal perception and streaming speech

Figure 2. Thinker-Talker architecture for multimodal perception and streaming speech. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

Qwen2.5-Omniの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5-。

Qwen2.5-Omniの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Gemma 3: Long-Context Multimodality with Efficient Interleaved Attention

Gemma 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3、128K。

arXiv GitHub HuggingFace

Gemma Team
Released: 2025-03-12

Architecture figure: The Gemma 3 report contains examples and analysis charts, but no architecture diagram.

ℹ️ 詳細情報

Gemma 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3、1B、4B、12B、27B、400M、896×896、256。

Gemma 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3、2、1,024-、32K、128K。

Gemma 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2T、4T、12T、14T。

Aya Vision: Multilingual Multimodality through Cross-Modal Model Merging

Aya Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv Website HuggingFace

Cohere Labs
Released: 2025-03-04

Architecture figure: The Aya Vision paper contains data and evaluation figures, but no model architecture diagram.

ℹ️ 詳細情報

Aya Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、32B、23。

Aya Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Phi-4-multimodal: Text, Vision, and Speech through Mixture-of-LoRAs

Phi-4-multimodalの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4-。

arXiv HuggingFace

Microsoft Phi Team
Released: 2025-03-03

Phi-4-multimodal: Text, Vision, and Speech through Mixture-of-LoRAs architecture: Vision and audio encoders, projectors, and modality-specific LoRA routes

Figure 1. Vision and audio encoders, projectors, and modality-specific LoRA routes. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

Phi-4-multimodalの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4-、5.6-、460。

Phi-4-multimodalの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 200K、128K。

Phi-4-multimodalの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3.5、4。

SigLIP 2: Multilingual Vision-Language Encoders with Native-Aspect-Ratio Support

SigLIP 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。

arXiv GitHub HuggingFace

Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muhammad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, et al.
Released: 2025-02-20

SigLIP 2: Multilingual Vision-Language Encoders with Native-Aspect-Ratio Support architecture: Combined contrastive, captioning, masked-prediction, and self-distillation objectives

Figure 1. Combined contrastive, captioning, masked-prediction, and self-distillation objectives. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

SigLIP 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。

SigLIP 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

SigLIP 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 10、12、109、90、2。

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

EVEv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

GitHub HuggingFace
EVEv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models architecture: EVEv2.0 architecture: lossless patch embeddings and text tokens enter a unified decoder-only VLM whose attention, feed-forward, and normalization layers use modality-specific weights.

Figure 3. EVEv2.0 architecture: lossless patch embeddings and text tokens enter a unified decoder-only VLM whose attention, feed-forward, and normalization layers use modality-specific weights. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

EVEv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、10M。

Qwen2.5-VL: Enhanced Vision-Language Capabilities in the Qwen Series

Qwen2.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5-。

arXiv GitHub HuggingFace
Qwen2.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Qwen2.5-VL: Enhanced Vision-Language Capabilities in the Qwen Series architecture: Qwen2.5-VL framework: a vision encoder processes native-resolution images and dynamic-FPS video into variable-length tokens for a Qwen2.5 language-model decoder with multimodal rotary position encoding.

Figure 1. Qwen2.5-VL framework: a vision encoder processes native-resolution images and dynamic-FPS video into variable-length tokens for a Qwen2.5 language-model decoder with multimodal rotary position encoding. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

Qwen2.5-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5-、3B、7B、72B、18。

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

VideoLLaMA 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace
VideoLLaMA 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding architecture: VideoLLaMA 3 pipeline with any-resolution vision tokenization and a differential frame pruner that compresses redundant video tokens before language-model processing.

Figure 3. VideoLLaMA 3 pipeline with any-resolution vision tokenization and a differential frame pruner that compresses redundant video tokens before language-model processing. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

VideoLLaMA 3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、1、2、3、4、1-。

UI-TARS: Pioneering Automated GUI Interaction with Native Agents

UI-TARSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 10。

arXiv GitHub HuggingFace
UI-TARSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

UI-TARS: Pioneering Automated GUI Interaction with Native Agents architecture: UI-TARS architecture and capability overview, connecting visual GUI observations and interaction histories to perception, action, system-level reasoning, and learning from prior experience.

Figure 4. UI-TARS architecture and capability overview, connecting visual GUI observations and interaction histories to perception, action, system-level reasoning, and learning from prior experience. Source paper, PDF p. 14. Figure notice.

ℹ️ 詳細情報

UI-TARSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1、2、3、4、2-、7B、72B。

MiniMax-01: Scaling Foundation Models with Lightning Attention

MiniMax-01の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 01、3.5-、4。

arXiv GitHub HuggingFace
MiniMax-01の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

MiniMax-01: Scaling Foundation Models with Lightning Attention architecture: MiniMax-Text-01 backbone architecture, interleaving Lightning Attention and softmax-attention transformer blocks with routed mixture-of-experts feed-forward layers.

Figure 3. MiniMax-Text-01 backbone architecture, interleaving Lightning Attention and softmax-attention transformer blocks with routed mixture-of-experts feed-forward layers. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

MiniMax-01の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 01、456、45.9、32、2、4。 MiniMax-01の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 01、694、100、2。

MiniCPM-o-2.6: A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming

MiniCPM-o-2.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.6、8B。

arXiv GitHub HuggingFace
MiniCPM-o-2.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

MiniCPM-o-2.6: A GPT-4o Level MLLM for Vision, Speech and Multimodal Live Streaming architecture: MiniCPM-o 2.6 end-to-end omni-modal streaming architecture, using time-division multiplexing to combine visual, audio, and query streams in a shared backbone with streaming speech decoding.

Official architecture diagram. MiniCPM-o 2.6 end-to-end omni-modal streaming architecture, using time-division multiplexing to combine visual, audio, and query streams in a shared backbone with streaming speech decoding. Primary source. Figure notice.

ℹ️ 詳細情報

MiniCPM-o-2.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.6、400M、300M、200M、5-7B。

MiniCPM-o-2.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連参照: https://arxiv.org/abs/2604.27393https://github.com/OpenBMB/MiniCPM-V。 値: 4.5、2026。

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

Eagle 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。

arXiv GitHub HuggingFace
Eagle 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models architecture: Eagle 2's tiled mixture of vision encoders, combining SigLIP and ConvNeXt features through dynamic image splitting, feature concatenation, pixel shuffle, and an MLP connector to the LLM.

Figure 11. Eagle 2's tiled mixture of vision encoders, combining SigLIP and ConvNeXt features through dynamic image splitting, feature concatenation, pixel shuffle, and an MLP connector to the LLM. Source paper, PDF p. 7. Figure notice.

ℹ️ 詳細情報

Eagle 2 adopts a “diversity first, then quality” data strategy, beginning with a large, diverse pool of over 180 data sources, followed by rigorous filtering and selection. The architecture uses a tiled mixture of vision encoders (MoVE), specifically SigLIP and ConvNeXt-XXLarge, with image tiling to handle high resolutions. Each image tile is encoded by channel-concatenated MoVE. The vision encoder outputs are concatenated and aligned with the LLM (Qwen2.5) via a simple MLP connector. A three-stage training recipe is used: Stage 1 trains the connector to align modalities; Stage 1.5 trains the full model on a large, diverse dataset; and Stage 2 fine-tunes on a high-quality instruction-tuning dataset. Crucially, all available visual instruction data is used in Stage 1.5, not just captioning/knowledge data. Balanced data packing addresses limitations in existing open-source frameworks. The core contribution is the detailed data strategy. This involves: (1) Data Collection: Building a highly diverse data pool (180+ sources) through both passive gathering (monitoring arXiv and Hugging Face) and proactive searching (addressing “bucket effect” via error analysis). (2) Data Filtering: Removing low-quality samples based on criteria like mismatched question-answer pairs, irrelevant image-question pairs, repeated text, and numeric formatting issues. (3) Data Selection: Choosing optimal subsets based on data source diversity, distribution, and K-means clustering on SSCD image embeddings to ensure balance across types (especially useful for chart data, etc.). (4) Data Augmentation: Mining information from input images through techniques like Chain-of-Thought (CoT) explanation generation, rule-based QA generation, and expanding short answers into longer ones. (5) Data Formating: remove unnecessary decorations. Training uses a three-stage approach: Eagle 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1。 Eagle 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5、21.6M。 Eagle 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、4.6M。

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Sa2VAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。

arXiv GitHub HuggingFace
Sa2VAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos architecture: Sa2VA model: text, prompts, images, and videos are encoded for an LLM, whose segmentation token is combined with SAM 2 features to decode image or video masks.

Figure 2. Sa2VA model: text, prompts, images, and videos are encoded for an LLM, whose segmentation token is combined with SAM 2 features to decode image or video masks. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

Sa2VAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、91、93、72,000、2,000、1.5、665K、17K、22K、214K、100K、3.5K、0.6K、1.7K、37K。

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

VideoChat-Flashの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace
VideoChat-Flashの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling architecture: VideoChat-Flash framework with hierarchical video-token compression: shared encoders and connectors first compress clips, then the LLM performs video-level compression for long-context inference.

Figure 3. VideoChat-Flash framework with hierarchical video-token compression: shared encoders and connectors first compress clips, then the LLM performs video-level compression for long-context inference. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

VideoChat-Flashの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、224、1、0.5M、2、3.5M、2.5M、3、1.1M、1.7M、0.7M、4、448、25、15、114,228、3,444,849。

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference

OmniVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 729、81、5-0.5B、400M。

arXiv HuggingFace
OmniVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference architecture: OmniVLM architecture: a vision transformer feeds a reshape-based projector that compresses image tokens before they join text tokens in the Qwen2.5-0.5B-Instruct language model.

Figure 1. OmniVLM architecture: a vision transformer feeds a reshape-based projector that compresses image tokens before they join text tokens in the Qwen2.5-0.5B-Instruct language model. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

OmniVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 91、729、93、81、1、400M、5-0.5B、2、3、81-、9.、1.。

Apollo: An Exploration of Video Understanding in Large Multimodal Models

Apolloの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3B、7B。

arXiv GitHub
Apolloの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Apollo: An Exploration of Video Understanding in Large Multimodal Models architecture: Apollo architecture: image and video encoders process N-frame clips, interpolated features are concatenated channel-wise, and a Perceiver resampler produces a fixed token set for the language model.

Figure 8. Apollo architecture: image and video encoders process N-frame clips, interpolated features are concatenated channel-wise, and a Perceiver resampler produces a fixed token set for the language model. Source paper, PDF p. 21. Figure notice.

ℹ️ 詳細情報

Apolloの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、1.5B、3B、7B、32、3-、8-32。 or clips is sufficient for efficient token integration. Training Stages is also disscussed, concluding that progressively unfreezing the different components in different stages leads to superior model training dynamics. Finally, training the Video Encoder is discussed. The paper concludes that Finetuning video encoders on only video data further improves overall performance, Apolloの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 Apolloの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

DeepSeek-VL2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace
DeepSeek-VL2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding architecture: DeepSeek-VL2 architecture: dynamic image tiling feeds a vision encoder and vision-language adapter whose image tokens join text tokens in a DeepSeek-MoE language model.

Figure 2. DeepSeek-VL2 architecture: dynamic image tiling feeds a vision encoder and vision-language adapter whose image tokens join text tokens in a DeepSeek-MoE language model. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

DeepSeek-VL2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1、2、3、384、9、729、1152、196、14、210、1.0B、2.8B、4.5B、1.2M、70、30、12M、800B。

Maya: An Instruction Finetuned Multilingual Multimodal Model

Mayaの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace
Mayaの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Maya: An Instruction Finetuned Multilingual Multimodal Model architecture: Maya's LLaVA-derived architecture, projecting multilingual SigLIP vision features into the embedding space of a multilingual language model for instruction following.

Figure 7. Maya's LLaVA-derived architecture, projecting multilingual SigLIP vision features into the embedding space of a multilingual language model for instruction following. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

Mayaの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連コード: Zv = g(Xv)WHv。 値: 1.5、23、8B、2-、35B、7,531、3、150K、10、2、558,000、7B、13B。

InternVL 2.5: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

InternVL 2.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.5、2.0、3.5-。

arXiv GitHub HuggingFace
InternVL 2.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

InternVL 2.5: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling architecture: InternVL 2.5's ViT-MLP-LLM architecture with pixel-unshuffle visual-token compression.

Figure 2. InternVL 2.5's ViT-MLP-LLM architecture with pixel-unshuffle visual-token compression. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

InternVL 2.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連コード: nmaxr。 値: 2.5、6B、300M、2-、1024、256、2.0、1、2、3。

PaliGemma 2: A Family of Versatile VLMs for Transfer

PaliGemma 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、3B、10B、28B。

arXiv GitHub HuggingFace
PaliGemma 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

PaliGemma 2: A Family of Versatile VLMs for Transfer architecture: PaliGemma 2 processes variable-resolution images with SigLIP, a linear projector, and Gemma 2.

Figure 1. PaliGemma 2 processes variable-resolution images with SigLIP, a linear projector, and Gemma 2. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

PaliGemma 2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、2B、9B、27B、3B、10B、28B、1、50、10、3。

ShowUI: Vision-Language-Action Modeling for GUI Agents

ShowUIの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2B。

arXiv GitHub HuggingFace

Kevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang, Shiwei Wu, Zechen Bai, Weixian Lei, Lijuan Wang, Mike Zheng Shou
Released: 2024-11-26

ShowUI: Vision-Language-Action Modeling for GUI Agents architecture: UI-guided visual-token selection and interleaved action history

Figure 3. UI-guided visual-token selection and interleaved action history. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

ShowUIの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2B、33、1.4。

ShowUIの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 256K。

ShowUIの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

SmolVLM: A Small, Efficient, and Open-Source Vision-Language Model

SmolVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2B、2.0。

arXiv GitHub HuggingFace
SmolVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

SmolVLM: A Small, Efficient, and Open-Source Vision-Language Model architecture: SmolVLM's SigLIP vision encoder, aggressive pixel-shuffle compression, and SmolLM2 language backbone.

Official architecture diagram. SmolVLM's SigLIP vision encoder, aggressive pixel-shuffle compression, and SmolLM2 language backbone. Primary source. Figure notice.

ℹ️ 詳細情報

SmolVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3.1、8B、1.7B、9、1.。

SmolVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連参照: https://huggingface.co/blog/smolvlm2。 値: 20、2025、256M、500M、2.2B。

AIMv2: Multimodal Autoregressive Pre-training of Large Vision Encoders

AIMv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace
AIMv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

AIMv2: Multimodal Autoregressive Pre-training of Large Vision Encoders architecture: AIMv2's prefix-attention vision encoder and joint autoregressive multimodal decoder.

Figure 1. AIMv2's prefix-attention vision encoder and joint autoregressive multimodal decoder. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

AIMv2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2B、300、3。

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

LLaVA-CoTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace
LLaVA-CoTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step architecture: LLaVA-CoT's Best-of-N, stage-wise beam-search, and stage-wise retracing inference procedures.

Figure 4. LLaVA-CoT's Best-of-N, stage-wise beam-search, and stage-wise retracing inference procedures. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

LLaVA-CoTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3.2-。

LLM2CLIP: Powerful Language Model Unlocks Richer Visual Representation

LLM2CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace
LLM2CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

LLM2CLIP: Powerful Language Model Unlocks Richer Visual Representation architecture: LLM2CLIP fine-tunes an LLM for caption discrimination before using it to train stronger CLIP representations.

Figure 1. LLM2CLIP fine-tunes an LLM for caption discrimination before using it to train stronger CLIP representations. Source paper, PDF p. 1. Figure notice.

ℹ️ 詳細情報

LLM2CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1、3、8B、2。 LLM2CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 15M、3M、103、12M、1B。

Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Tarsier2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5。

arXiv GitHub HuggingFace
Tarsier2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Architecture figure: The paper contains capability, dataset, DPO-construction, and benchmark figures but inherits Qwen2-VL and provides no architecture overview; a training diagram would be misleading.

ℹ️ 詳細情報

Tarsier2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 11、40、150K、15、585K。

Janus and Janus-Pro: Decoupled Visual Understanding and Generation

Janus and Janus-Proの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Janus Janus Pro GitHub HuggingFace
DeepSeek-AI
Released: 2024-10-17

Janus and Janus-Pro: Decoupled Visual Understanding and Generation architecture: Janus-Pro decouples visual understanding and generation encoders around a shared autoregressive transformer.

Figure 3. Janus-Pro decouples visual understanding and generation encoders around a shared autoregressive transformer. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

Janus and Janus-Proの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Janus and Janus-Proの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2025、1.5B、7B、90M、72M。

ARIA: An Open Multimodal Native Mixture-of-Experts Model

ARIAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4-。

arXiv GitHub HuggingFace
ARIAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Architecture figure: ARIA’s visual encoder, projection layer, and fine-grained MoE are described in prose and a configuration table; its only numbered model figure visualizes expert specialization, not architecture.

ℹ️ 詳細情報

ARIAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3.5B、3.9B、24.9B、66、2、6、438M、4-、1、6.4T、8K、400B、3、64K、4、20B。 ARIAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Emu3: Next-Token Prediction across Text, Image, and Video

Emu3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B。

arXiv GitHub HuggingFace

Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo, Quan Sun, Yufeng Cui, Jinsheng Wang, Fan Zhang, Yueze Wang, Zhen Li, Qiying Yu, Yingli Zhao, Yulong Ao, Xuebin Min, Tao Li, Boya Wu, et al.
Released: 2024-09-27

Emu3: Next-Token Prediction across Text, Image, and Video architecture: A single next-token objective across text, image, and video tokens

Figure 1. A single next-token objective across text, image, and video tokens. Source paper, PDF p. 1. Figure notice.

ℹ️ 詳細情報

Emu3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Emu3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、0.2、0.8。

Emu3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Molmo and PixMo: Open Weights, Open Data, and Grounded Pointing

Molmo and PixMoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace

Matt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi, Yue Yang, Jae Sung Park, Mohammadreza Salehi, Niklas Muennighoff, Kyle Lo, Luca Soldaini, Jiasen Lu, et al.
Released: 2024-09-25

Molmo and PixMo: Open Weights, Open Data, and Grounded Pointing architecture: Vision encoding, multimodal connection, language decoding, and grounded pointing

Figure 2. Vision encoding, multimodal connection, language decoding, and grounded pointing. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

Molmo and PixMoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14、1B、7B、72B。

Molmo and PixMoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Molmo and PixMoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Llama 3.2-Vision: Enhanced Multimodal Capabilities Built on Llama 3

Llama 3.2-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3.2-、3、11B、90B。

GitHub HuggingFace
Llama 3.2-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Architecture figure: The official source describes the image encoder and cross-attention adapter in prose; benchmark tables and marketing artwork are not legitimate substitutes.

ℹ️ 詳細情報

Llama 3.2-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連コード: <|image|>brave_searchwolfram_alpha。 値: 3.2-、3、6B、2023、3.2。

NVLM: Open Frontier-Class Multimodal LLMs

NVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.0、1-。

arXiv GitHub HuggingFace
NVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

NVLM: Open Frontier-Class Multimodal LLMs architecture: NVLM-X, NVLM-H, and NVLM-D share a dynamic-high-resolution visual pathway but integrate vision differently.

Figure 3. NVLM-X, NVLM-H, and NVLM-D share a dynamic-high-resolution visual pathway but integrate vision differently. Source paper, PDF p. 9. Figure notice.

ℹ️ 詳細情報

NVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連コード: <tile_1>。 値: 6B、5、6、1024、256、72B、2-、34B、1-、115M、40、40-。

Pixtral 12B: A Cutting-Edge Open Multimodal Language Model

Pixtral 12Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 12B、12-。

arXiv GitHub HuggingFace
Pixtral 12Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Pixtral 12B: A Cutting-Edge Open Multimodal Language Model architecture: Pixtral combines a variable-resolution vision encoder with a 128K-context multimodal decoder.

Figure 3. Pixtral combines a variable-resolution vision encoder with a 128K-context multimodal decoder. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

Pixtral 12Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 12B、128K、91、93、12-。

Pixtral 12Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連参照: https://mistral.ai/news/pixtral-largehttps://mistral.ai/news/mistral-small-4/。 値: 2024、4、2026。

VILA-U: Fully Autoregressive Visual Understanding and Generation

VILA-Uの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub

VILA-U Team
Released: 2024-09-06

VILA-U: Fully Autoregressive Visual Understanding and Generation architecture: A shared visual tokenizer, autoregressive model, and modality decoders

Figure 1. A shared visual tokenizer, autoregressive model, and modality decoders. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

VILA-Uの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

VILA-Uの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Qwen2-VL: A Powerful Open-Source Vision-Language Model for Image and Video Understanding

Qwen2-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

GitHub HuggingFace
Bai, Jinze and Bai, Shuai and Yang, Shusheng and Wang, Shijie and Tan, Sinan and Wang, Peng and Lin, Junyang and Zhou, Chang and Zhou, Jingren

Qwen2-VL: A Powerful Open-Source Vision-Language Model for Image and Video Understanding architecture: Qwen2-VL's M-RoPE decomposes multimodal position encoding into temporal, height, and width components.

Figure 3. Qwen2-VL's M-RoPE decomposes multimodal position encoding into temporal, height, and width components. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

Qwen2-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 600M、20。

EAGLE: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

EAGLEの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace
Min Shi, Fuxiao Liu, Shihao Wang, Shijia Liao, Subhashree Radhakrishnan, De-An Huang, Hongxu Yin, Karan Sapra, Yaser Yacoob, Humphrey Shi, Bryan Catanzaro, Andrew Tao, Jan Kautz, Zhiding Yu, Guilin Liu

EAGLE: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders architecture: EAGLE explores mixtures of vision experts and alternative fusion strategies for multimodal language models.

Figure 2. EAGLE explores mixtures of vision experts and alternative fusion strategies for multimodal language models. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

EAGLEの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Show-o: Autoregressive Language and Discrete-Diffusion Vision in One Transformer

Show-oの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace

Show Lab
Released: 2024-08-22

Show-o: Autoregressive Language and Discrete-Diffusion Vision in One Transformer architecture: Causal text attention and full-attention discrete image diffusion

Figure 2. Causal text attention and full-attention discrete image diffusion. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

Show-oの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Show-oの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Idefics3-8B: Building and Better Understanding Vision-Language Models

Idefics3-8Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B。

arXiv HuggingFace
Hugo Laurençon, Andrés Marafioti, Victor Sanh, Léo Tronchon

Idefics3-8B: Building and Better Understanding Vision-Language Models architecture: Idefics3 maps vision-encoder features into interleaved visual tokens consumed by an autoregressive language model.

Figure 1. Idefics3 maps vision-encoder features into interleaved visual tokens consumed by an autoregressive language model. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

Idefics3-8Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、3.1、1.5、4、169、364、1820、13.7-。

Transfusion: Next-Token Text Prediction and Continuous Image Diffusion

Transfusionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv

Meta FAIR
Released: 2024-08-20

Transfusion: Next-Token Text Prediction and Continuous Image Diffusion architecture: A shared Transformer with autoregressive text and continuous image diffusion

Figure 1. A shared Transformer with autoregressive text and continuous image diffusion. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

Transfusionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Transfusionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、16。

mPLUG-Owl3: Hyper-Attention for Long Image Sequences

mPLUG-Owl3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace

mPLUG-Owl Team
Released: 2024-08-09

mPLUG-Owl3: Hyper-Attention for Long Image Sequences architecture: Vision encoding and Hyper-Attention blocks inside the language model

Figure 2. Vision encoding and Hyper-Attention blocks inside the language model. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

mPLUG-Owl3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

mPLUG-Owl3の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

VITA: Towards Open-Source Interactive Omni Multimodal LLM

VITAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace
VITAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

VITA: Towards Open-Source Interactive Omni Multimodal LLM architecture: VITA unifies text, audio, image, and video inputs with state tokens, an LLM, and speech output.

Figure 2. VITA unifies text, audio, image, and video inputs with state tokens, an LLM, and speech output. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

VITAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 300M、256、24-、25、2。

LLaVA-OneVision: Easy Visual Task Transfer

LLaVA-OneVisionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv Website HuggingFace
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, Chunyuan Li

LLaVA-OneVision: Easy Visual Task Transfer architecture: LLaVA-OneVision extends the minimal LLaVA vision-encoder, projector, and LLM architecture to multiple visual signals.

Figure 1. LLaVA-OneVision extends the minimal LLaVA vision-encoder, projector, and LLM architecture to multiple visual signals. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

LLaVA-OneVisionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、2-。

LLaVA-OneVisionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連参照: https://arxiv.org/abs/2509.23661https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-1.5。 値: 28、2025、1.5、85M、26M、64B、16,000。

VILA²: VILA Augmented VILA

VILA²の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv HuggingFace
Yunhao Fang, Ligeng Zhu, Yao Lu, Yan Wang, Pavlo Molchanov, Jang Hyun Cho, Marco Pavone, Song Han, Hongxu Yin

VILA²: VILA Augmented VILA architecture: VILA² improves itself through generic model-in-the-loop recaptioning and specialist-model augmentation.

Figure 1. VILA² improves itself through generic model-in-the-loop recaptioning and specialist-model augmentation. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

VILA²の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、8B、34B、6B。

INF-LLaVA: High-Resolution Image Perception for Multimodal Large Language Models

INF-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace
Yiwei Ma, Zhibin Wang, Xiaoshuai Sun, Weihuang Lin, Qiang Zhou, Jiayi Ji, Rongrong Ji

INF-LLaVA: High-Resolution Image Perception for Multimodal Large Language Models architecture: INF-LLaVA combines dual-perspective cropping, CLIP encoding, feature recombination, enhancement, and language reasoning.

Figure 2. INF-LLaVA combines dual-perspective cropping, CLIP encoding, feature recombination, enhancement, and language reasoning. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

INF-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

SlowFast-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv HuggingFace
Mingze Xu, Mingfei Gao, Zhe Gan, Hong-You Chen, Zhengfeng Lai, Haiming Gang, Kai Kang, Afshin Dehghan

SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models architecture: SlowFast-LLaVA combines low-frame-rate spatial detail with high-frame-rate motion features.

Figure 2. SlowFast-LLaVA combines low-frame-rate spatial detail with high-frame-rate motion features. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

SlowFast-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8、64。

EVLM: An Efficient Vision-Language Model for Visual Understanding

EVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv HuggingFace
Kaibing Chen, Dong Shen, Hanwen Zhong, Huasong Zhong, Kui Xia, Di Xu, Wei Yuan, Yifei Hu, Bin Wen, Tianke Zhang, Changyi Liu, Dewen Fan, Huihui Xiao, Jiahong Wu, Fan Yang, Size Li, Di Zhang

EVLM: An Efficient Vision-Language Model for Visual Understanding architecture: EVLM injects hierarchical EVA2-CLIP features into the language model through gated cross-attention layers.

Figure 2. EVLM injects hierarchical EVA2-CLIP features into the language model through gated cross-attention layers. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

EVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.4B、8、40、16、14B、1.0。

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

InternLM-XComposer-2.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.5、7B。

arXiv GitHub HuggingFace
Pan Zhang, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Rui Qian, Lin Chen, Qipeng Guo, Haodong Duan, Bin Wang, Linke Ouyang, Songyang Zhang, Wenwei Zhang, Yining Li, Yang Gao, Peng Sun, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Hang Yan, Conghui He, Xingcheng Zhang, Kai Chen, Jifeng Dai, Yu Qiao, Dahua Lin, Jiaqi Wang

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output architecture: InternLM-XComposer-2.5's framework supports text, single and multiple images, and video inputs.

Figure 5. InternLM-XComposer-2.5's framework supports text, single and multiple images, and video inputs. Source paper, PDF p. 6. Figure notice.

ℹ️ 詳細情報

InternLM-XComposer-2.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.5、2、2-、14、7B。

OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

OMG-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace
Tao Zhang, Xiangtai Li, Hao Fei, Haobo Yuan, Shengqiong Wu, Shunping Ji, Chen Change Loy, Shuicheng Yan

OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding architecture: OMG-LLaVA connects OMG-Seg visual tokens and prompts to an LLM that can decode segmentation outputs.

Figure 3. OMG-LLaVA connects OMG-Seg visual tokens and prompts to an LLM that can decode segmentation outputs. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

OMG-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Cambrian-1: Vision-Centric Multimodal LLMs

Cambrian-1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1、8B、13B、34B。

arXiv GitHub

Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Ziteng Wang, Rob Fergus, Yann LeCun, Saining Xie
Released: 2024-06-24

Cambrian-1: Vision-Centric Multimodal LLMs architecture: Spatial Vision Aggregation across multiple visual encoders and decoder layers

Figure 8. Spatial Vision Aggregation across multiple visual encoders and decoder layers. Source paper, PDF p. 13. Figure notice.

ℹ️ 詳細情報

Cambrian-1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1、3、1.5、2-。

Cambrian-1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.5M、7M。

Cambrian-1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7M、10M。

EVE: Unveiling Encoder-Free Vision-Language Models

EVEの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 35M。

arXiv GitHub HuggingFace
EVEの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

EVE: Unveiling Encoder-Free Vision-Language Models architecture: EVE combines patch embedding, a decoder-only backbone, patch alignment, and next-word prediction without a separate vision encoder.

Figure 2. EVE combines patch embedding, a decoder-only backbone, patch alignment, and next-word prediction without a separate vision encoder. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

EVEの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連コード: <CLS><SPL>。 値: 7B、14、16M、33M、665K、1。 EVEの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 33M、1.5、665K。

Ovis: Structural Visual-Text Embedding Alignment

Ovisの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv Ovis 2.5 GitHub HuggingFace

Alibaba International Digital Commerce
Released: 2024-06-14

Ovis: Structural Visual-Text Embedding Alignment architecture: Visual-token probability distributions and structural visual embedding lookup

Figure 3. Visual-token probability distributions and structural visual embedding lookup. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

Ovisの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、1.6。

Ovisの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 26、2025、5、1B、34B、15、2。

Parrot: Multilingual Visual Instruction Tuning

Parrotの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
Hai-Long Sun, Da-Wei Zhou, Yang Li, Shiyin Lu, Chao Yi, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, De-Chuan Zhan, Han-Jia Ye

Parrot: Multilingual Visual Instruction Tuning architecture: PARROT aligns multilingual visual features through a projector, multilingual mixture of experts, and language model.

Figure 5. PARROT aligns multilingual visual features through a projector, multilingual mixture of experts, and language model. Source paper, PDF p. 6. Figure notice.

ℹ️ 詳細情報

Parrotの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14、5-。

ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

ConvLLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace
Chunjiang Ge, Sijie Cheng, Ziming Wang, Jiale Yuan, Yuan Gao, Jun Song, Shiji Song, Gao Huang, Bo Zheng

ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models architecture: ConvLLaVA uses a hierarchical ConvNeXt vision encoder to compress visual tokens between stages.

Figure 1. ConvLLaVA uses a hierarchical ConvNeXt vision encoder to compress visual tokens between stages. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

ConvLLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 576。

Phi-3-Vision and Phi-3.5-Vision: Compact Long-Context Multimodal Reasoning

Phi-3-Vision and Phi-3.5-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3-、4.2B、3.5-。

Technical Report Official Release HuggingFace

Microsoft Phi Team
Released: 2024-05-21

Architecture figure: The cited Phi-3 report contains no Phi-3-Vision architecture diagram.

ℹ️ 詳細情報

Phi-3-Vision and Phi-3.5-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3-、3、128K、4.2B。

Phi-3-Vision and Phi-3.5-Visionの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3-、21、2024、3.5-、22。

CogVLM2: Enhanced Vision-Language Models for Image and Video Understanding

CogVLM2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace
Wenyi Hong, Weihan Wang, Ming Ding, Wenmeng Yu, Qingsong Lv, Yan Wang, Yean Cheng, Shiyu Huang, Junhui Ji, Zhao Xue, Lei Zhao, Zhuoyi Yang, Xiaotao Gu, Xiaohan Zhang, Guanyu Feng, Da Yin, Zihan Wang, Ji Qi, Xixuan Song, Peng Zhang, Debing Liu, Bin Xu, Juanzi Li, Yuxiao Dong, Jie Tang

CogVLM2: Enhanced Vision-Language Models for Image and Video Understanding architecture: CogVLM2 processes high-resolution images and video frames through a ViT encoder, adapter, and visual-language decoder.

Figure 2. CogVLM2 processes high-resolution images and video frames through a ViT encoder, adapter, and visual-language decoder. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

CogVLM2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Chameleon: Mixed-Modal Early-Fusion Foundation Models

Chameleonの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub HuggingFace

Chameleon Team
Released: 2024-05-16

Chameleon: Mixed-Modal Early-Fusion Foundation Models architecture: Mixed-modal early-fusion tokenization and autoregressive generation

Figure 1. Mixed-modal early-fusion tokenization and autoregressive generation. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

Chameleonの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、34B。

Chameleonの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4.4。

Chameleonの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

PaliGemma: A Versatile and Transferable 3B Vision-Language Model

PaliGemmaの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2B、40。

arXiv GitHub HuggingFace
Lucas Beyer, Andreas Steiner, André Susano Pinto, Alexander Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, Thomas Unterthiner, Daniel Keysers, Skanda Koppula, Fangyu Liu, Adam Grycner, Alexey Gritsenko, Neil Houlsby, Manoj Kumar, Keran Rong, Julian Eisenschlos, Rishabh Kabra, Matthias Bauer, Matko Bošnjak, Xi Chen, Matthias Minderer, Paul Voigtlaender, Ioana Bica, Ivana Balazevic, Joan Puigcerver, Pinelopi Papalampidi, Olivier Henaff, Xi Xiong, Radu Soricut, Jeremiah Harmsen, Xiaohua Zhai

PaliGemma: A Versatile and Transferable 3B Vision-Language Model architecture: PaliGemma connects a SigLIP image encoder to a Gemma autoregressive decoder language model.

Figure 1. PaliGemma connects a SigLIP image encoder to a Gemma autoregressive decoder language model. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

PaliGemmaの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3-、2B、30。

xGen-MM (BLIP-3): An Open-Source Framework for Building Powerful and Responsible Large Multimodal Models

xGen-MM (BLIP-3)の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3。

arXiv HuggingFace
Le Xue, Manli Shu, Anas Awadalla, Jun Wang, An Yan, Senthil Purushwalkam, Honglu Zhou, Viraj Prabhu, Yutong Dai, Michael S Ryoo, Shrikant Kendre, Jieyu Zhang, Can Qin, Shu Zhang, Chia-Chih Chen, Ning Yu, Juntao Tan, Tulika Manoj Awalgaonkar, Shelby Heinecke, Huan Wang, Yejin Choi, Ludwig Schmidt, Zeyuan Chen, Silvio Savarese, Juan Carlos Niebles, Caiming Xiong, Ran Xu

xGen-MM (BLIP-3): An Open-Source Framework for Building Powerful and Responsible Large Multimodal Models architecture: BLIP-3 feeds interleaved visual and text tokens through a scalable vision-token sampler into a pretrained language model.

Figure 2. BLIP-3 feeds interleaved visual and text tokens through a scalable vision-token sampler into a pretrained language model. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

xGen-MM (BLIP-3)の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3、2。

MANTIS: Mastering Multi-Image Understanding Through Interleaved Instruction Tuning

MANTISの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub Gradio
Dongfu Jiang, Xuan He, Huaye Zeng, Cong Wei, Max Ku, Qian Liu, Wenhu Chen

Architecture figure: The paper contains capability examples, dataset statistics and illustrations, and case studies, none of which is a model architecture or system diagram.

ℹ️ 詳細情報

MANTISの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 721K、3、8B。

Moondream-next: Compact Vision-Language Model with Enhanced Capabilities

Moondream-nextの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.9B。

arXiv GitHub HuggingFace

Architecture figure: The rolling model card and official repository provide implementation details but no legitimate architecture or system figure.

ℹ️ 詳細情報

Moondream-nextの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.9B。

Idefics2

Idefics2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、7B。

arXiv Gradio
Idefics2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Idefics2 architecture: Idefics2 maps vision-encoder features into interleaved visual tokens consumed by an autoregressive language model.

Figure 2. Idefics2 maps vision-encoder features into interleaved visual tokens consumed by an autoregressive language model. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

Idefics2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、7B、2.5。

InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

InternLM-XComposer2-4KHDの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 336、4K。

arXiv
Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Songyang Zhang, Haodong Duan, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Zhe Chen, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Kai Chen, Conghui He, Xingcheng Zhang, Jifeng Dai, Yu Qiao, Dahua Lin, Jiaqi Wang

InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD architecture: InternLM-XComposer2-4KHD dynamically partitions high-resolution images into local patches while retaining a global thumbnail.

Figure 4. InternLM-XComposer2-4KHD dynamically partitions high-resolution images into local patches while retaining a global thumbnail. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

InternLM-XComposer2-4KHDの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4K、336。

MM1: Methods, Analysis, and Insights from Multimodal Pre-training

MM1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv Project

Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Floris Weers, Anton Belyi, Haotian Zhang, et al.
Released: 2024-03-14

Architecture figure: MM1 presents an architecture design space and ablations, but no definitive final-model diagram.

ℹ️ 詳細情報

MM1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3B、30B。

MM1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

MM1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

DeepSeek-VL: Towards Real-World Vision-Language Understanding

DeepSeek-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub

Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, Yaofeng Sun, Chengqi Deng, Hanwei Xu, Zhenda Xie, Chong Ruan

DeepSeek-VL: Towards Real-World Vision-Language Understanding architecture: DeepSeek-VL trains its hybrid vision encoder, vision-language adaptor, and language model across three progressive stages.

Figure 3. DeepSeek-VL trains its hybrid vision encoder, vision-language adaptor, and language model across three progressive stages. Source paper, PDF p. 12. Figure notice.

ℹ️ 詳細情報

DeepSeek-VL: Employs a hybrid vision encoder architecture, fusing a SigLIP-L encoder for semantic understanding with a SAM-B encoder for high-resolution detail extraction. This allows for efficient processing of 1024x1024 images while capturing both global and fine-grained visual features. A two-layer hybrid MLP adapter then integrates these features with the DeepSeek LLM backbone. The model is pre-trained on a diverse dataset encompassing web screenshots, PDFs, OCR, charts, and knowledge-based content from sources like Common Crawl, Web Code, E-books, and arXiv articles. This pretraining is further refined using a curated instruction-tuning dataset based on real user scenarios and categorized into a comprehensive taxonomy covering recognition, conversion, analysis, reasoning, evaluation, and safety tasks. By combining this diverse data with its unique architecture and fusion strategies, DeepSeek-VL aims to deliver robust performance across a wide range of real-world vision-language applications.

AnyGPT: Unified Any-to-Any Multimodal Modeling with Discrete Tokens

AnyGPTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub

OpenMOSS
Released: 2024-02-19

AnyGPT: Unified Any-to-Any Multimodal Modeling with Discrete Tokens architecture: Unified discrete sequence modeling across speech, text, image, and music

Figure 1. Unified discrete sequence modeling across speech, text, image, and music. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

AnyGPTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

AnyGPTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 108,000-。

SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

SPHINX-Xの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub Model
SPHINX-Xの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models architecture: SPHINX-X combines mixed visual experts, padded-tile skip tokens, high-resolution partitioning, and unified training.

Figure 3. SPHINX-X combines mixed visual experts, padded-tile skip tokens, high-resolution partitioning, and unified training. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

SPHINX-Xの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

LLaVA 1.6: LLaVA-NeXT Improved reasoning, OCR, and world knowledge

LLaVA 1.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5。

GitHub
LLaVA 1.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

LLaVA 1.6: LLaVA-NeXT Improved reasoning, OCR, and world knowledge architecture: LLaVA-NeXT's AnyRes scheme partitions high-resolution images into a configurable grid of local views.

Official architecture diagram. LLaVA-NeXT's AnyRes scheme partitions high-resolution images into a configurable grid of local views. Primary source. Figure notice.

ℹ️ 詳細情報

LLaVA 1.6の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5、1、32。

MiniCPM-V: A GPT-4V Level MLLM on Your Phone

MiniCPM-Vの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2.5、3。

arXiv GitHub HuggingFace
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, Qianyu Chen, Huarong Zhou, Zhensheng Zou, Haoye Zhang, Shengding Hu, Zhi Zheng, Jie Zhou, Jie Cai, Xu Han, Guoyang Zeng, Dahai Li, Zhiyuan Liu, Maosong Sun

MiniCPM-V: A GPT-4V Level MLLM on Your Phone architecture: MiniCPM-V combines a visual encoder, shared compression layer, language model, and adaptive high-resolution encoding.

Figure 3. MiniCPM-V combines a visual encoder, shared compression layer, language model, and adaptive high-resolution encoding. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

MiniCPM-Vの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14、1024、64、96、2B、8B、2.5。

MiniCPM-Vの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 関連参照: https://arxiv.org/abs/2509.18154https://huggingface.co/openbmb/MiniCPM-V-4.6https://github.com/OpenBMB/MiniCPM-V。 値: 4.5、2025、4.6、2026、1B、400M、5-0.8B。

MouSi: Poly-Visual-Expert Vision-Language Models

MouSiの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
MouSiの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

MouSi: Poly-Visual-Expert Vision-Language Models architecture: MouSi integrates heterogeneous visual experts through a fusion network and projector into a language model.

Figure 2. MouSi integrates heterogeneous visual experts through a fusion network and projector into a language model. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

MouSiの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 5、558K。

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

InternLM-XComposer2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。

arXiv GitHub Gradio
InternLM-XComposer2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model architecture: InternLM-XComposer2 applies Partial-LoRA only to visual tokens while preserving language-only behavior.

Figure 2. InternLM-XComposer2 applies Partial-LoRA only to visual tokens while preserving language-only behavior. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

InternLM-XComposer2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

MoE-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub Gradio
MoE-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models architecture: MoE-LLaVA routes multimodal tokens through sparse experts added to a vision-encoder, projector, and language-model backbone.

Figure 2. MoE-LLaVA routes multimodal tokens through sparse experts added to a vision-encoder, projector, and language-model backbone. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

MoE-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

moondream1 and moondream2

moondream1 and moondream2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5。

GitHub Gradio
@vikhyatk

Architecture figure: Official repository and model cards describe the implementation but provide no legitimate family architecture or system figure.

ℹ️ 詳細情報

moondream1 and moondream2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.6B、1.5、1.86B。

FireLLaVA

FireLLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 34B。

Model

Architecture figure: A generic LLaVA diagram would not document FireLLaVA’s own contribution and should not be substituted.

ℹ️ 詳細情報

FireLLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 34B、588K。

COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training

COSMOの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv Website

COSMOの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training architecture: CosMo handles image and video inputs through a language model trained with contrastive and language-modeling objectives.

Figure 3. CosMo handles image and video inputs through a language model trained with contrastive and language-modeling objectives. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

COSMOの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 128、14。

TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

TinyGPT-Vの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。

arXiv GitHub Gradio
TinyGPT-Vの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones architecture: TinyGPT-V projects frozen visual-backbone and Q-Former outputs through two linear layers into Phi-2.

Figure 4. TinyGPT-V projects frozen visual-backbone and Q-Former outputs through two linear layers into Phi-2. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

TinyGPT-Vの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、2.7。

MobileVLM: A Fast, Strong and Open Vision Language Assistant for Mobile Devices

MobileVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14。

arXiv GitHub
MobileVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

MobileVLM: A Fast, Strong and Open Vision Language Assistant for Mobile Devices architecture: MobileVLM connects a visual encoder to MobileLLaMA through a lightweight downsample projector.

Figure 1. MobileVLM connects a visual encoder to MobileLLaMA through a lightweight downsample projector. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

MobileVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14。

Alpha-CLIP: A CLIP Model Focusing on Wherever You Want

Alpha-CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub

Alpha-CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Alpha-CLIP: A CLIP Model Focusing on Wherever You Want architecture: Alpha-CLIP extends CLIP with an alpha channel that focuses visual encoding on a specified region.

Figure 3. Alpha-CLIP extends CLIP with an alpha channel that focuses visual encoding on a specified region. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

Alpha-CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 Alpha-CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 400M、5B。

Nous-Hermes-2-Vision - Mistral 7B

Nous-Hermes-2-Vision - Mistral 7Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2-、2.5、400M。

Model
Nous-Hermes-2-Vision - Mistral 7Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Architecture figure: The card describes SigLIP integration and training data in prose but contains no legitimate architecture or system figure.

ℹ️ 詳細情報

Nous-Hermes-2-Vision - Mistral 7Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2-、2.5-、7B、400M、3B、220K、60K、150K、50K、2.5。

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

SPHINXの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。

arXiv GitHub SPHINXの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models architecture: SPHINX jointly mixes tuning tasks, visual embeddings, and model weights in one multimodal architecture.

Figure 3. SPHINX jointly mixes tuning tasks, visual embeddings, and model weights in one multimodal architecture. Source paper, PDF p. 6. Figure notice.

ℹ️ 詳細情報

SPHINXの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、400M。

Florence-2: A Deep Dive into its Unified Architecture and Multi-Task Capabilities

Florence-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。

arXiv HuggingFace
Bin Xiao, Haiping Wu, Weijian Xu, Xiyang Dai, Houdong Hu, Yumao Lu, Michael Zeng, Ce Liu, Lu Yuan

Florence-2: A Deep Dive into its Unified Architecture and Multi-Task Capabilities architecture: Florence-2 combines an image encoder and multimodality encoder-decoder through a unified task-prompt interface.

Figure 2. Florence-2 combines an image encoder and multimodality encoder-decoder through a unified task-prompt interface. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

Florence-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。

u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model

u-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
u-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model architecture: u-LLaVA unifies modality alignment with task-specific projectors, decoders, and patched downstream modules.

Figure 1. u-LLaVA unifies modality alignment with task-specific projectors, decoders, and patched downstream modules. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

u-LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14、58K、23K。

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

LLaVA-Plusの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
LLaVA-Plusの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents architecture: LLaVA-Plus retrieves skills, invokes tools, consumes their results, and generates a final response.

Figure 2. LLaVA-Plus retrieves skills, invokes tools, consumes their results, and generates a final response. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

LLaVA-Plusの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

OtterHD: A High-Resolution Multi-modality Model

OtterHDの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B。

arXiv GitHub

OtterHDの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Architecture figure: The paper’s figures cover demonstrations, benchmark construction, throughput, resolution, and loss; none is a legitimate architecture substitute.

ℹ️ 詳細情報

OtterHDの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B、2。

CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding

CoVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv
CoVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding architecture: CoVLM vision module and communication-token framework.

Figure 2. CoVLM vision module and communication-token framework. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

CoVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 97。

GLaMM: Pixel Grounding Large Multimodal Model

GLaMMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
Hanoona Rasheed, Muhammad Maaz, Sahal Shaji Mullappilly, Abdelrahman Shaker, Salman Khan, Hisham Cholakkal, Rao M. Anwer, Erix Xing, Ming-Hsuan Yang, Fahad S. Khan

GLaMM: Pixel Grounding Large Multimodal Model architecture: GLaMM architecture for scene-, region-, and pixel-level grounding.

Figure 2. GLaMM architecture for scene-, region-, and pixel-level grounding. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

GLaMMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7.5、810。

Fuyu-8B: A Multimodal Architecture for AI Agents

Fuyu-8Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B。

Link Model
Fuyu-8Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Fuyu-8B: A Multimodal Architecture for AI Agents architecture: Fuyu-8B projects image patches directly into a decoder-only Transformer.

Official architecture diagram. Fuyu-8B projects image patches directly into a decoder-only Transformer. Primary source. Figure notice.

ℹ️ 詳細情報

Fuyu-8Bの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 8B。

PaLI-3 Vision Language Models: Smaller, Faster, Stronger

PaLI-3 Vision Language Modelsの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3、2B、3B。

arXiv GitHub
PaLI-3 Vision Language Modelsの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

PaLI-3 Vision Language Models: Smaller, Faster, Stronger architecture: PaLI-3 connects a contrastively pretrained SigLIP encoder to an encoder-decoder UL2 Transformer.

Figure 1. PaLI-3 connects a contrastively pretrained SigLIP encoder to an encoder-decoder UL2 Transformer. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

PaLI-3 Vision Language Modelsの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 3、2B、3B。

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

MiniGPT-v2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、2-。

arXiv
MiniGPT-v2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning architecture: MiniGPT-v2 architecture with frozen ViT, token concatenation, projection, and LLaMA-2.

Figure 2. MiniGPT-v2 architecture with frozen ViT, token concatenation, projection, and LLaMA-2. Source paper, PDF p. 3. Figure notice.


ℹ️ 詳細情報

MiniGPT-v2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2-、7-、20M。

BakLLaVA

BakLLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、1.5、2、13B。

GitHub Model

Architecture figure: The official model card and repository describe a LLaVA-on-Mistral derivative but publish no model-specific architecture figure.

ℹ️ 詳細情報

BakLLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 7B、1.5、2、13B、600K、150K、558K、158K、450K、40K。

Ferret: Refer and Ground Anything Anywhere at Any Granularity

Ferretの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
Ferretの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Ferret: Refer and Ground Anything Anywhere at Any Granularity architecture: Ferret hybrid region representation, spatial-aware sampler, and complete model architecture.

Figure 3. Ferret hybrid region representation, spatial-aware sampler, and complete model architecture. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

Ferretの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.1。

LLaVA 1.5: Improved Baselines with Visual Instruction Tuning

LLaVA 1.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5。

arXiv
LLaVA 1.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

LLaVA 1.5: Improved Baselines with Visual Instruction Tuning architecture: LLaVA-1.5-HD grid-based encoding for arbitrary image resolutions.

Figure 2. LLaVA-1.5-HD grid-based encoding for arbitrary image resolutions. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

LLaVA 1.5の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5、11。

CogVLM: Visual Expert for Pretrained Language Models

CogVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
CogVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

CogVLM: Visual Expert for Pretrained Language Models architecture: CogVLM input pathway and visual-expert Transformer block.

Figure 4. CogVLM input pathway and visual-expert Transformer block. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

CogVLMの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1.5、40、2B、700M。

MetaCLIP: Demystifying CLIP Data

MetaCLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
MetaCLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Architecture figure: The paper contributes data curation rather than a model architecture; Figure 5 is a data-pipeline case study.

ℹ️ 詳細情報

MetaCLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 400。

Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Qwen-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub

Qwen-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond architecture: Qwen-VL visual-language architecture across pretraining, multitask pretraining, and supervised finetuning.

Figure 3. Qwen-VL visual-language architecture across pretraining, multitask pretraining, and supervised finetuning. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

Qwen-VLの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

IDEFICS

IDEFICSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 80B、4。

Model

Architecture figure: The official launch post and model card have capability and performance illustrations but no IDEFICS-specific architecture diagram.

ℹ️ 詳細情報

IDEFICSの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 80、4。

BLIVA: A Simple Multimodal LLM for Better Handling of Text-rich Visual Questions

BLIVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
BLIVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

BLIVA: A Simple Multimodal LLM for Better Handling of Text-rich Visual Questions architecture: BLIVA architecture with frozen image encoder, Q-Former, patch projection, and frozen LLM.

Figure 2. BLIVA architecture with frozen image encoder, Q-Former, patch projection, and frozen LLM. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

BLIVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

KOSMOS-2: Grounding Multimodal Large Language Models to the World

KOSMOS-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、1。

arXiv GitHub Gradio
KOSMOS-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

KOSMOS-2: Grounding Multimodal Large Language Models to the World architecture: KOSMOS-2 system overview for multimodal grounding and referring.

Figure 1. KOSMOS-2 system overview for multimodal grounding and referring. Source paper, PDF p. 1. Figure notice.

ℹ️ 詳細情報

KOSMOS-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、1、256。

LaVIN: Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models

LaVINの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
LaVINの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

LaVIN: Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models architecture: LaVIN architecture and Mixture-of-Modality Adaptation mechanism.

Figure 2. LaVIN architecture and Mixture-of-Modality Adaptation mechanism. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

LaVINの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

InstructBLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。

arXiv GitHub Gradio
InstructBLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning architecture: InstructBLIP architecture with instruction-aware Q-Former and frozen LLM.

Figure 3. InstructBLIP architecture with instruction-aware Q-Former and frozen LLM. Source paper, PDF p. 5. Figure notice.

ℹ️ 詳細情報

InstructBLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2、26、11。

ImageBind: One Embedding Space To Bind Them All

ImageBindの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
ImageBindの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

ImageBind: One Embedding Space To Bind Them All architecture: ImageBind aligns six modalities in one shared embedding space.

Figure 2. ImageBind aligns six modalities in one shared embedding space. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

ImageBindの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

LLaVA: Large Language and Vision Assistant - Visual Instruction Tuning

LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub

LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

LLaVA: Large Language and Vision Assistant - Visual Instruction Tuning architecture: LLaVA connects CLIP visual features to a language model through a learned projection.

Figure 1. LLaVA connects CLIP visual features to a language model through a learned projection. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

LLaVAの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4、158K。

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

MiniGPT-4の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4。

arXiv GitHub
MiniGPT-4の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models architecture: MiniGPT-4 architecture with ViT, Q-Former, linear projection, and Vicuna.

Figure 1. MiniGPT-4 architecture with ViT, Q-Former, linear projection, and Vicuna. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

MiniGPT-4の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4、20,000、256、3,500。

SigLIP: Sigmoid Loss for Language Image Pre-Training

SigLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas Beyer

Architecture figure: The paper changes the training loss, not the encoder architecture; Figure 1 is a distributed loss-implementation mock-up.

ℹ️ 詳細情報

SigLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

OpenFlamingoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14、7B。

arXiv GitHub
OpenFlamingoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models architecture: OpenFlamingo-9B interleaved image-and-text system interface.

Figure 2. OpenFlamingo-9B interleaved image-and-text system interface. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

OpenFlamingoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 14、7-、7B、2B、64。

PaLM-E: An Embodied Multimodal Language Model

PaLM-Eの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
PaLM-Eの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

PaLM-E: An Embodied Multimodal Language Model architecture: PaLM-E combines sensor encoders and a language model for embodied and visual-language tasks.

Figure 1. PaLM-E combines sensor encoders and a language model for embodied and visual-language tasks. Source paper, PDF p. 1. Figure notice.

ℹ️ 詳細情報

PaLM-Eの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

KOSMOS-1: Language Is Not All You Need: Aligning Perception with Language Models

KOSMOS-1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1。

arXiv GitHub
KOSMOS-1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

KOSMOS-1: Language Is Not All You Need: Aligning Perception with Language Models architecture: KOSMOS-1 multimodal input, embedding, language-model, and output overview.

Figure 1. KOSMOS-1 multimodal input, embedding, language-model, and output overview. Source paper, PDF p. 1. Figure notice.

ℹ️ 詳細情報

KOSMOS-1の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 1、2B、400M、700M。

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

BLIP-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。

arXiv GitHub Gradio
BLIP-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models architecture: BLIP-2 bridges a frozen image encoder and frozen LLM through a two-stage Q-Former.

Figure 1. BLIP-2 bridges a frozen image encoder and frozen LLM through a two-stage Q-Former. Source paper, PDF p. 1. Figure notice.

ℹ️ 詳細情報

BLIP-2の構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 2。

MULTIINSTRUCT: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning

MULTIINSTRUCTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
MULTIINSTRUCTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Architecture figure: Figures cover examples, task taxonomy, performance, and attention; the paper publishes no model-specific architecture figure.

ℹ️ 詳細情報

MULTIINSTRUCTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

PaLI: A Jointly-Scaled Multilingual Language-Image Model

PaLIの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
PaLIの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

PaLI: A Jointly-Scaled Multilingual Language-Image Model architecture: PaLI combines a scalable ViT with an encoder-decoder Transformer.

Figure 1. PaLI combines a scalable ViT with an encoder-decoder Transformer. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

PaLIの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 4、10、100、17B。

Flamingo: a Visual Language Model for Few-Shot Learning

Flamingoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv
Flamingoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

Flamingo: a Visual Language Model for Few-Shot Learning architecture: Flamingo architecture for interleaved visual inputs and free-form text output.

Figure 3. Flamingo architecture for interleaved visual inputs and free-form text output. Source paper, PDF p. 4. Figure notice.

ℹ️ 詳細情報

Flamingoの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

BLIP: Bootstrapping Language-Image Pre-training

BLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
Junnan Li, Dongxu Li, Caiming Xiong, Steven Hoi

BLIP: Bootstrapping Language-Image Pre-training architecture: BLIP multimodal mixture-of-encoder-decoder architecture and training objectives.

Figure 2. BLIP multimodal mixture-of-encoder-decoder architecture and training objectives. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

BLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 12M。

GLIP: Grounded Language-Image Pre-training

GLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
GLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

GLIP: Grounded Language-Image Pre-training architecture: GLIP image and language encoders with deep fusion and word-region alignment.

Figure 2. GLIP image and language encoders with deep fusion and word-region alignment. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

GLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

FROZEN: Multimodal Few-Shot Learning with Frozen Language Models

FROZENの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 50。

arXiv
FROZENの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

FROZEN: Multimodal Few-Shot Learning with Frozen Language Models architecture: FROZEN trains a vision encoder through a frozen language model.

Figure 2. FROZEN trains a vision encoder through a frozen language model. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

FROZENの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 50。

CLIP: Contrastive Language-Image Pre-training

CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 400。

arXiv GitHub
CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

CLIP: Contrastive Language-Image Pre-training architecture: CLIP dual-encoder contrastive training and zero-shot classification approach.

Figure 1. CLIP dual-encoder contrastive training and zero-shot classification approach. Source paper, PDF p. 2. Figure notice.

ℹ️ 詳細情報

CLIPの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 400。

ViT: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

ViTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

arXiv GitHub
ViTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。

ViT: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale architecture: ViT patch embedding, Transformer encoder, and classification-token architecture.

Figure 1. ViT patch embedding, Transformer encoder, and classification-token architecture. Source paper, PDF p. 3. Figure notice.

ℹ️ 詳細情報

ViTの構造、学習方法、データ、モダリティ統合、設計上の特徴に関する要約です。 値: 300M、100。

重要な参考資料