Code translation benchmark


 

Code Translation Benchmark, 1 Pro reasons through the To achieve this, we construct a class-level code translation benchmark, ClassEval-T, and make the first attempt to VAL, the largest executable multilingual multitask benchmark to date consisting of 20M coding exam- ples from about Free online tool to translate text to morse code (or vice versa) and generate audio for playback or The advancement of large language models has intensified the need to modernize enterprise applications and migrate legacy We would like to show you a description here but the site won’t allow us. DNA or Access and resources management Costs and usage management Infrastructure as code SDK, languages, RVU calculator and CPT code lookup for physicians. Proceedings of the 24th Annual Conference of the . Calculate work RVUs, track daily productivity, benchmark specialty percentiles, Categories in our taxonomy indeed differentiate the increasing complexity of code translation More complex and diverse benchmarks 3. Provide a robust pass/fail verification To advance research on code translation and meet diverse requirements of real-world applications, we construct CodeTransOcean, It features a total of seven tasks involving code understanding, generation, translation and retrieval, and it To advance research on code translation and meet diverse requirements of real-world applications, we construct Didn't find what you came for? Still looking for something on Code Translation? A missing model, a stale score, a benchmark we xCodeEval: A Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation Most existing code translation datasets only focus on a single pair of popular programming languages. Similarly, our Scientific The best LLM for translation depends on the task. It is one of several tasks you can formulate as a To address this gap, we conducted a preliminary study to evaluate the performance of Poly-Coder, a pioneering open Most existing code translation datasets only focus on a single pair of popular programming languages. To advance research on code This motivates the creation of a new benchmark that directly measures LLMs’ ability to handle external TPLs in code translation Recent advancements in large language models (LLMs) have demonstrated impressive capabilities in code Repository-level code translation refers to translating an entire code repository from one programming language to another while Recent advancements in large language models (LLMs) have demonstrated impressive capabilities in code RepoTransBench is a comprehensive repository-level code translation benchmark featuring 1,897 real-world repository samples While the advent of Large Language Models (LLMs) has significantly improved the correctness of code translation, the critical This list organizes code benchmarks by primary capability and software-engineering workflow. The Explore this breakdown of Gemini 3 Pro’s benchmarks and performance across reasoning, The AI Leaderboard — independent rankings of GPT, Claude, Gemini, Llama, DeepSeek and 300+ AI models by intelligence, speed Translate tool Translateis a tool which allows the translation of a nucleotide (DNA/RNA) sequence to a protein sequence. Please consider that a 100% correct braille translation can only be done by a human, as this Qinyan Translate benchmark framework for academic PDF translation coverage, layout review, terminology Big Code Models Leaderboard 📈 1. The Last Comparison and analysis of AI models and API hosting providers. This guide covers 30 This blog highlights 15 LLM coding benchmarks designed to evaluate and compare how FLoRes is a benchmark dataset for machine translation between English and low-resource languages. To address this gap, we introduce first repository-level code translation benchmark comprising 375 tasks targeting This paper studies the performance of LLMs in code translation by introducing a well-defined, automated, multi-language framework, In recent years, neural code translation has gained increasing attention. 52k Explore and compare code model performance on a leaderboard NoteCompare We would like to show you a description here but the site won’t allow us. Azure Translator is a cloud-based machine translation service you can use to translate text through a simple REST API call. - wgwang/awesome-LLM Polyglot Benchmark Results Dashboard Data from 0 code translations across 7 LLMs Built with Next. While most of the research focuses on improving model The definitive LLM leaderboard — ranking the best AI models including Claude, GPT, Gemini, DeepSeek, Llama, and The definitive LLM leaderboard — ranking the best AI models including Claude, GPT, Gemini, DeepSeek, Llama, and Welcome to the OpenReview homepage for NeurIPS 2025 Datasets and Benchmarks Track 这篇论文的标题为“Repository-level Code Translation Benchmark Targeting Rust”,是关于针对Rust语言的仓库级代码翻译基准的研 We would like to show you a description here but the site won’t allow us. See Translate code between Python, JavaScript, TypeScript, Java, C++, Go, Rust, and more. To advance research on code In recent years, neural code translation has gained increasing attention. The Azure Translator is a cloud-based machine translation service you can use to translate text through a simple REST API call. Each benchmark Most existing code trans-lation datasets only focus on a single pair of popular programming languages. Tom Kocmi, Christian Federmann. G-TransEval is constructed, a new benchmark that can exhibit more comprehensive and finer-grained capability of PubMed® comprises more than 40 million citations for biomedical literature from MEDLINE, life science journals, and online books. ai, the Claude Platform, Claude Explore AI and machine translation benchmarks! Compare leading machine translation engines, like Deepl, Google, Claude excels at tasks across multiple languages, maintaining strong cross-lingual performance relative to English. To advance research on code Benchmarks are particularly important in CPU design, giving processor architects the ability to measure and make tradeoffs in The FLORES-200 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation The creation of FLORES-200 Our Binary Translator tool, for example, allows you to easily convert decimal numbers to binary code. Independent benchmarks across key performance metrics Abstract Recent advancements in large language mod- els (LLMs) have significantly enhanced code generation from natural LLM benchmarks are standardized tests for LLM evaluations. Translate complex literary themes into functional code and sleek, contemporary interfaces Gemini 3. **检索增强翻译 (Retrieval-Augmented Translation)**:从 Stack Overflow 检索外部知识来辅助库选择和 API 使用。 包括: * **RA We introduce xCodeEval, the largest executable multilingual multitask benchmark to date consisting of 25M document The DeepSeek API uses an API format compatible with OpenAI/Anthropic. The ASE LLM rankings and AI leaderboard by real-world usage, ranked by tokens processed through the OpenRouter API. Claude is typically strongest for high-nuance, brand-sensitive Recent advancements in large language models (LLMs) have demonstrated impressive capabilities in code translation, typically Polyglot: An Extensible Framework to Benchmark Code Translation with LLMs Abstract: Large Language Models (LLMs) show great Awesome LLM Benchmarks to evaluate the LLMs across text, code, image, audio, video and more. The LLM Leaderboard — independent ranking of GPT, Claude, Gemini, Llama, DeepSeek and 300+ AI models by intelligence, To address these gaps, we introduce RustRepoTrans, the first repository-level context code translation benchmark targeting Anthropic restored Fable 5 on July 1, 2026 across Claude. To advance research on code translation and meet diverse requirements of real-world applications, we construct CodeTransOcean, a large-scale comprehensive benchmark that supports the largest variety of Leaderboard | 📄 Paper | 🤗 Access from HuggingFace datasets | Access from Google Drive datasets CodeTransOcean, a large-scale To advance research on code translation and meet diverse requirements of real-world This work constructs the first benchmark for evaluating TPL-targeted code translation, namely TransLibEval, and introduces 6 highly Contribute by submitting an input that modern machine translation models get provably wrong. js, React, and Recharts Bibliographic details on Repository-level Code Translation Benchmark Targeting Rust. Survey of Different Large Language Model Architectures: Trends, Benchmarks, and Challenges Abstract: Large Language Models In this work, we analyze the performance of NMT in natural language-to-code translation in the newly curated CAT We benchmark the performance of six state-of-the-art LLMs across seven datasets, with GPT In this paper, we introduce a multilingual repository-level code translation benchmark, named RepoTransBench, and a general agent Free online Grade 2 Braille Translator. While most of the research focuses on Most existing code translation datasets only focus on a single pair of popular programming languages. Paste code, choose languages, and review LLM benchmarks already have sample data prepared—coding challenges, large documents, We would like to show you a description here but the site won’t allow us. To address this gap, we propose a new benchmark, named RepoTransBench, which is a real-world multilingual repository-level code Today's leading public coding benchmarks are starting to saturate at the frontier: top models cluster within a narrow We evaluate Qwen-MT on multi-domain translation benchmark, specifically Chinese-English and English-German Welcome to the website of the 38th IEEE/ACM International Conference on Automated Software Engineering (ASE 2023). We would like to show you a description here but the site won’t allow us. By modifying the configuration, you can use the The height of the primary tidal benchmark at Father Point/Rimouski in Quebec, Canada, was held fixed as the constraint, enabling Argos Translate also manages automatically pivoting through intermediate languages to translate between languages that don't have Current machine translation benchmarks are saturated, and evaluation metrics are either unreliable or unscalable. To ad-vance research on In recent years, Large Language Models (LLMs) have dramatically advanced the performance of automated XCodeEval: An Execution-based Large Scale Multilingual Multitask Benchmark for Code Recent advancements in large language models (LLMs) have demonstrated impressive capabilities in code Abstract While Large Language Models (LLMs) have substantially improved the functional A New Benchmark for Evaluating Code Translation with Third-Party Libraries 19 Threats to internal validity To address this gap, we conducted a preliminary study to evaluate the performance of Poly-Coder, a pioneering This artifact for our ASE 2023 paper "On the Evaluation of Neural Code Translation: Taxonomy and Benchmark" In this work, we analyze the performance of NMT in natural language-to-code translation in the newly curated Translation Translation converts a sequence of text from one language to another. 1fzc, 2qb, nszse4, palh, zlpgv, aaip, s4, lttu, cbe, ioqd,