SenseTime Research

Advancing native multimodal AI to explore the frontier of intelligence and shape new AI paradigms. Built on rigorous academic foundations and a forward-looking research vision, SenseTime Research turns frontier science into productive capabilities for industries.

Technology Updates

Multimodal

SenseTime Releases SenseNova 6.7 Flash-Lite, Cutting Token Consumption by 60%

SenseTime introduced SenseNova 6.7 Flash-Lite, a new lightweight multimodal agent model, alongside the SenseNova Token Plan and open-sourced SenseNova-Skills on GitHub.

Read Original
Foundation Model

SenseNova U1 Open-Sourced, Moving Toward Unified Understanding and Generation

SenseNova U1 is a native unified model for understanding and generation, built on SenseTime's NEO-unify architecture to bring multimodal understanding, reasoning, and generation into one model.

Read Original
Multimodal

72x Faster Inference and 7-Minute Long Video Generation with Kairos 3.0-4B

Daxiao Robot open-sourced Kairos 3.0-4B, an embodied-native world model that integrates multimodal understanding, generation, and prediction to advance embodied AI.

Read Original

Research Articles

Research progress across foundation models, multimodal AI, reasoning, agents, and coding models.
Multimodal

Simulating the Physical World! SenseTime and NTU S-Lab Introduce Spatial-Reasoner, a 3D Spatial Reasoning Multimodal Large Language Model

SenseTime and NTU S-Lab introduce Spatial-Reasoner, a 3D spatial reasoning multimodal large language model for physical-world understanding and embodied AI.

Reasoning

SenseNova-SI-1.3 Advances Spatial Intelligence Across Eight Benchmarks

SenseTime open-sourced SenseNova-SI-1.3, improving spatial measurement, viewpoint transformation, spatial reasoning, and question answering.

Reasoning

SenseNova-MARS Breaks Through the Ceiling of Multimodal Search Reasoning

SenseNova-MARS combines search, perception, and reasoning to improve information grounding and answer generation for complex multimodal tasks.

Multimodal

SekoTalk Brings Real-Time Voice-Driven Digital Humans Closer to Production

SekoTalk advances real-time interactive digital humans through speech driving, expression synchronization, and low-latency generation.

Foundation Model

From Data Fusion to Native Architecture: SenseTime Introduces NEO

Developed with NTU S-Lab, NEO lays a new architectural foundation for the SenseNova multimodal model family.

Reasoning

SenseNova Open-Source Models Achieve Breakthroughs in Spatial Intelligence

SenseNova-SI models deliver stronger spatial understanding and reasoning, with leading performance among open-source models at comparable scale.

Reasoning

SenseNova Open-Source Models Achieve Breakthroughs in Spatial Intelligence

SenseNova-SI models deliver stronger spatial understanding and reasoning, with leading performance among open-source models at comparable scale.

Multimodal

SenseTime's Perspective: Why We Are Firmly Committed to Multimodal General Intelligence

SenseTime co-founder, executive director, and chief scientist Dahua Lin explains the logic, technical path, practice, and future direction behind multimodal general intelligence.

Multimodal

SenseNova 6.5 Upgrades AI from a Tool to More Human-Like Interaction

SenseNova 6.5 strengthens multimodal understanding, deep reasoning, and interleaved text-image thinking, moving AI toward more natural interaction.

Foundation Model

2025 SenseTime Scholarship Announced: 30 AI Rising Stars Unveiled

The 2025 SenseTime Scholarship recognizes 30 young AI researchers and continues to support foundational AI research and frontier applications.

Agent

SenseTime Launches Wuneng Embodied AI Platform for Self-Evolving AI in the Physical World

Built around the Kaiwu world model, Wuneng helps robots and intelligent devices explore, learn, and evolve in the physical world.

Multimodal

SenseNova V6 Wins Two Championships in Language and Multimodal Benchmarks

SenseNova V6 achieved leading results in both general language and multimodal capability evaluations, demonstrating strong reasoning and multimodal understanding.