LongCat-AudioDiT

A non-autoregressive diffusion text-to-speech model operating directly in waveform latent space, with 1B and 3.5B variants achieving SOTA zero-shot voice cloning.

Independent third-party resource. Not affiliated with or endorsed by LongCat or Meituan.

Overview

LongCat-AudioDiT is documented in the meituan-longcat GitHub repository and is listed in the publisher API documentation. The model is the same entity referenced in both sources.

Documented capabilities

The statements below are described in the cited primary documentation. They describe what the cited documentation says the model is, not an independent performance evaluation.

  • Zero-shot voice cloning: Seed-ZH SIM 0.818, Seed-Hard SIM 0.797.
  • Two model sizes: 1B and 3.5B parameters.
  • Wav-VAE encoder with >2,000x compression (24kHz to 11.7Hz frame rate).
  • Adaptive Projection Guidance (APG) replacing classifier-free guidance.
  • Trained on 1M hours of Chinese and English speech.
  • Waveform latent space direct generation — no mel-spectrogram intermediates.

Access and license

Open source. Weights and code publicly released.

Independent-site notes

This profile is a source-based summary. It does not contain any benchmark, evaluation, or quality claim produced by this site.

No independent testing has been published by this site.

Any performance figures attributed to the model in third-party materials are vendor-reported and have not been independently verified on this site.

Sources

FAQ

Does LongCat-AudioDiT support voice cloning with a single reference sample?
Yes. LongCat-AudioDiT achieves zero-shot voice cloning, meaning it can clone a speaker's voice from a single reference audio sample without fine-tuning. The 3.5B variant reaches a speaker similarity of 0.818 on Seed-ZH.
Is this page an official LongCat or Meituan page?
No. This page is published by an independent third-party site. It is not affiliated with, endorsed by, or sponsored by LongCat or Meituan.
How often is this page updated?
This page was last verified on 2026-07-27. Content is reviewed when new publisher documentation or model releases become available.

Independent third-party disclosure

This page is published by an independent third-party site. It is not affiliated with, endorsed by, sponsored by, or operated by Meituan, LongCat, or any of their affiliates. The content summarizes publicly-available primary documentation and does not represent the views of any referenced organization.

Last reviewed: 2026-07-27