No matter where you start, build and scale your AI with ByteCompute.
All categories and models you can try out and seamlessly integrate in your projects

automatic-speech-recognition
A weakly supervised pre-trained version of the Whisper model, optimized for high-speed Automatic Speech Recognition (ASR) and speech translation. By significantly reducing the number of decoder layers to 4 while maintaining the robust large-v3 encoder architecture, this 'Turbo' variant offers an 8.8x speedup compared to large-v3 with minimal degradation in Word Error Rate (WER). It is specifically designed as a high-efficiency alternative for low-latency production environments.

TEXT
The flagship Mixture-of-Experts (MoE) model from the Qwen3.5 series, featuring 122B total and 10B active parameters. This unified vision-language foundation excels in multimodal reasoning, complex coding, and native 'thinking mode' tasks. Utilizing fine-grained FP8 quantization, it offers exceptional throughput and reduced VRAM footprint on H100/L40S GPUs, while supporting a massive 262K context window for long-horizon agentic applications.

VIDEO
A state-of-the-art Diffusion Transformer (DiT) foundation model with 22 billion parameters. Unlike traditional video models, LTX-2 is natively designed for synchronized audio-video generation within a single unified latent space. It excels at maintaining temporal consistency and high-fidelity motion, making it a powerful backend for creative AI pipelines that require seamless audiovisual coherence.

automatic-speech-recognition
The Whisper large-v3 is a pre-trained model for Automatic Speech Recognition (ASR) and speech translation. It features a robust Transformer encoder-decoder architecture designed for state-of-the-art accuracy across a wide range of languages and audio conditions.

IMAGE
The FLUX.2 [klein] model family are our fastest image models to date. FLUX.2 [klein] unifies generation and editing in a single compact architecture, delivering state-of-the-art quality with end-to-end inference in as low as under a second.

TEXT
gemma-4-31B-it is a 30.7-billion parameter dense multimodal model. It is built using the same research infrastructure as Gemini 3, offering frontier-level performance for reasoning, agentic workflows, and high-accuracy coding. It natively supports text and image inputs with a massive 256K context window.

TEXT
Qwen3.6-35B-A3B is an efficient mixture-of-experts (MoE) model with 35B total parameters and only 3B active. It delivers strong agentic coding performance, significantly outperforming Qwen3.5-35B-A3B and competing with larger dense models like Qwen3.5-27B and Gemma4-31B. Supporting both multimodal thinking and non-thinking modes, it is a highly versatile open-source model, now available on Qwen Studio, via API, and as open weights.

OCR
PaddleOCR Service is a synchronous HTTP API for extracting text and document structure from images. Each request uploads a single image, the server runs inference, and the result is returned in the same response — no polling required.

TEXT
Qwen3.8-27B-FP8 is a deployment-friendly 27B dense vision-language model with fine-grained FP8 quantization. It excels at coding, professional work, research, and long-horizon agentic tasks, supports text, image, and video inputs, flexible thinking control, and a 262K native context window.
Contact our sales team to discuss your enterprise needs and deployment options.
Get Started