LILT Launches AURORA, the First Multilingual AI Leaderboard That Measures Frontier Models on Non-English Enterprise Agentic Tasks Grounded in Language and Culture
SAN FRANCISCO, Sept. 30, 2026
Press Release Disclaimer: This is a press release distributed through the XPR Media network. It has not been independently verified by our newsroom.

![]()
LILT Launches AURORA, the First Multilingual AI Leaderboard That Measures Frontier Models on Non-English Enterprise Agentic Tasks Grounded in Language and Culture
PR Newswire
SAN FRANCISCO, Sept. 30, 2026
SAN FRANCISCO, Sept. 30, 2026 /PRNewswire/ — LILT today launched AURORA, the industry’s first Multilingual AI Leaderboard that measures frontier models on non-English agentic, multimodal, and socio-cultural tasks through scientifically rigorous evaluation.

Enterprises are deploying agents worldwide, yet nearly every published measure of frontier progress relies on English-centric or translated benchmarks. With agents doing customer-facing work, performance in every language now carries commercial weight.
“The industry is choosing models on a scoreboard that stops at English. That gap used to cost you an awkward translation. Now that agents are taking real actions in the real world, it costs you a wrong decision, and you find out from your users.” — Spence Green, CEO and co-founder, LILT
Measuring Multilingual Performance:
AURORA addresses the multilingual gap by testing models against LILT’s multilingual benchmark suite featuring tasks designed and verified by native-language domain experts. These tasks represent real-world enterprise applications, such as software development, customer support, and complex workflows, each grounded in language, region and culture.
Specifically, AURORA provides visibility across LILT’s multilingual benchmarks, including:
- Multilingual Terminal-bench: Coding tasks representing challenges in software written for native-language users.
- Multilingual τ³-bench: Multi-turn customer support in airline, telecom, retail, and banking.
- Multilingual MultiChallenge: Long-context instruction-following, memory and self-coherence.
- Multilingual GAIA-v2-LILT: Agentic reasoning and tool use.
AURORA is developed and managed by LILT’s Applied AI practice, a PhD-led research team with 10+ years’ experience in multilingual AI.
Their latest analysis shows model quality can differ substantially between languages. For example, in coding, GPT 5.5 performs best in Spanish, Claude Opus 5.5 wins in Japanese, while Muse Spark 1.3 leads in Serbian.
Availability
AURORA is live at https://aurora.lilt.com/, where readers can compare models by language and task. AI teams interested in private, custom benchmarks can contact LILT at contact@lilt.com.
About LILT
LILT is the leading agentic AI solution to make anything multilingual at scale for enterprise, public sector, and Frontier labs. LILT’s Applied AI services give AI teams private multilingual benchmarks, evaluation, and training data needed to ship better agents and models. Leading Frontier labs and organizations like NVIDIA, Intel, and L’Oréal rely on LILT to expand global reach. Learn more at lilt.com.
View original content to download multimedia:https://www.prnewswire.com/news-releases/lilt-launches-aurora-the-first-multilingual-ai-leaderboard-that-measures-frontier-models-on-non-english-enterprise-agentic-tasks-grounded-in-language-and-culture-302894528.html
SOURCE LILT

