LongCat-Flash-Omni
A 560B-parameter open-source omni-modal model with 27B activated, excelling at real-time audio-visual interaction — text, image, audio, and video in a single end-to-end framework.
Independent third-party resource. Not affiliated with or endorsed by LongCat or Meituan.
Overview
LongCat-Flash-Omni is documented in the meituan-longcat GitHub repository and is listed in the publisher API documentation. The model is the same entity referenced in both sources.
Documented capabilities
The statements below are described in the cited primary documentation. They describe what the cited documentation says the model is, not an independent performance evaluation.
- Real-time audio-visual interaction — text, image, audio, and video within a single framework.
- 560B total parameters (MoE), 27B activated per token on average.
- 128K token context window with multi-turn dialogue support.
- Open-source SOTA on Omni-Bench and WorldSense omni-modal benchmarks.
- Low-latency streaming multi-modal input/output with chunk-wise feature interleaving.
- Strong vision benchmarks: MMBench-EN 87.5, DocVQA 91.8, OCRBench 84.9.
- Audio capabilities: ASR (LibriSpeech CER 1.57), S2TT, TTS, and audio understanding.
- Video understanding: MVBench 75.2, VideoMME (w/ audio) 78.2.
- Integrates with LongCat-Audio-Codec (0.43–0.87 kbps, ~100ms latency).
- Deployed via LongCat App (iOS, Android) for real-time voice and image interaction.
Access and license
MIT License. Open-source release includes model weights and inference code.
Documented context window: 128,000 tokens (per the cited publisher API documentation).
Independent-site notes
This profile is a source-based summary. It does not contain any benchmark, evaluation, or quality claim produced by this site.
No independent testing has been published by this site.
Any performance figures attributed to the model in third-party materials are vendor-reported and have not been independently verified on this site.
Sources
- LongCat-Flash-Omni Technical Report (arXiv:2511.00279)
Primary sourcePublished 2025-11-03Accessed 2026-07-26
Technical report describing the omni-modal architecture, curriculum-inspired progressive training, Modality-Decoupled Parallelism training scheme sustaining >90% text-only throughput, and comprehensive benchmark results across omni-modal, vision, video, and audio tasks.
- LongCat-Flash-Omni GitHub Repository
Primary sourceAccessed 2026-07-26
Official repository under MIT License. Contains model architecture overview, evaluation results, and deployment instructions.
- LongCat-Flash-Omni on HuggingFace
Primary sourceAccessed 2026-07-26
Model weights and model card with comprehensive benchmark tables for omni-modal, vision, video, and audio evaluations.
- Meituan Tech Post — LongCat-Flash-Omni Announcement
Publisher documentationPublished 2025-11-03Accessed 2026-07-26
Publisher announcement detailing the launch of LongCat App alongside the model release, end-to-end evaluation with 250 users and 10 expert reviewers.
FAQ
- What is LongCat-2.0?
- LongCat-2.0 is a model documented in the meituan-longcat GitHub organization and listed in the publisher API documentation.
- Where can the primary source for LongCat-2.0 be found?
- The meituan-longcat GitHub repository is the primary source. The repository contains the model code, configuration, and license file.
- Does LongCat-2.0 have a documented API?
- Yes. The publisher API documentation lists LongCat-2.0 as a supported model on both OpenAI-compatible and Anthropic-compatible endpoints.
- Is this page an official LongCat or Meituan page?
- No. This page is published by an independent third-party site that summarizes publicly-available primary documentation. It is not affiliated with, endorsed by, or sponsored by LongCat or Meituan.
Independent third-party disclosure
This page is published by an independent third-party site. It is not affiliated with, endorsed by, sponsored by, or operated by Meituan, LongCat, or any of their affiliates. The content summarizes publicly-available primary documentation and does not represent the views of any referenced organization.
Last reviewed: 2026-07-26