{"id":3635,"date":"2026-09-23T10:14:27","date_gmt":"2026-09-23T10:14:27","guid":{"rendered":"https:\/\/dev95.site\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/"},"modified":"2026-09-23T10:14:27","modified_gmt":"2026-09-23T10:14:27","slug":"android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks","status":"publish","type":"post","link":"https:\/\/dev95.site\/ar\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/","title":{"rendered":"Android Bench 2.0: Pushing the frontier with challenging long-horizon tasks"},"content":{"rendered":"<div id=\"dev95-428736727\" class=\"dev95-- dev95-entity-placement\"><script async=\"async\" data-cfasync=\"false\" src=\"https:\/\/pl27862732.profitableratecpmnetwork.com\/2ad7a50e0bbc23ac6801d7b77c501463\/invoke.js\"><\/script>\r\n<div id=\"container-2ad7a50e0bbc23ac6801d7b77c501463\"><\/div><\/div><div><meta content=\"https:\/\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEhCs6gPNr-l6f79eAyix8OZ59gg6K5y8QVTb6vuU2mNR9qdIlN2VUvGzbTenI-pIEGhMYql-E-t7Hs2Z0vI_UYnHte1w3vPRpjk7E0DPenuSkt-3gUM3y5GYZKHgciA4o3Ox2oxVkNuHiCwUX1WKCUkQhzBAd2FJhFiB-k5UKXYA67hXQHRVdFwrjHWFLQ\/s2049\/Bench 2.0 Metadata-bench.png\" style=\"clear: right; float: right; margin-bottom: 1em; margin-left: 1em;\"><br \/>\n<img data-recalc-dims=\"1\" decoding=\"async\" src=\"https:\/\/i0.wp.com\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEhCs6gPNr-l6f79eAyix8OZ59gg6K5y8QVTb6vuU2mNR9qdIlN2VUvGzbTenI-pIEGhMYql-E-t7Hs2Z0vI_UYnHte1w3vPRpjk7E0DPenuSkt-3gUM3y5GYZKHgciA4o3Ox2oxVkNuHiCwUX1WKCUkQhzBAd2FJhFiB-k5UKXYA67hXQHRVdFwrjHWFLQ\/s2049\/Bench%202.0%20Metadata-bench.png?w=1280&#038;ssl=1\" style=\"display: none;\"><\/p>\n<div><\/div>\n<div style=\"margin-top: -12px;\">\n  <i>Posted by Matthew McCullough, VP, Product Management, Android Developer<\/i>\n<\/div>\n<div class=\"separator\" style=\"clear: both; text-align: center;\"><a href=\"https:\/\/i0.wp.com\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEjA-sKC09IqUbbrWx4LuwZLfnBPs2Z9Z29K2yvRXEb-cbZmbTFfoX9qg7LYLNsKZ3hXMim2O0BpwHPiwtX2IsG2QsUEbr0WflIwDa5iEi9gR4epKClpCxVs86yAFKs4EXP48sxmE7iFaWeDSCNq_48V8_GXB_OcH0v2LTzCsAKQ4GXp4qc9C-K205lKQxk\/s4292\/Android%20Bench%202.0%20Blogger-bench%20%282%29.png?ssl=1\" style=\"clear: left; float: left; margin-bottom: 1em; margin-right: 1em;\"><img data-recalc-dims=\"1\" decoding=\"async\" border=\"0\" data-original-height=\"1300\" data-original-width=\"4292\" src=\"https:\/\/i0.wp.com\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEjA-sKC09IqUbbrWx4LuwZLfnBPs2Z9Z29K2yvRXEb-cbZmbTFfoX9qg7LYLNsKZ3hXMim2O0BpwHPiwtX2IsG2QsUEbr0WflIwDa5iEi9gR4epKClpCxVs86yAFKs4EXP48sxmE7iFaWeDSCNq_48V8_GXB_OcH0v2LTzCsAKQ4GXp4qc9C-K205lKQxk\/s1600\/Android%20Bench%202.0%20Blogger-bench%20%282%29.png?w=1280&#038;ssl=1\"><\/a><\/div>\n<p><\/p>\n<div>\n<p><i><br \/><\/i><\/p>\n<p>When we first launched Android Bench, we built a rigorous foundation for evaluating how large language models (LLMs) assist developers with real-world Android tasks. As AI models and agents rapidly evolve, we\u2019ve been updating our methodology, such as aligning our benchmark framework with <a href=\"https:\/\/android-developers.googleblog.com\/2026\/07\/android-bench-llm-measurement.html\" target=\"_blank\">the Harbor framework<\/a>. Today <b>we\u2019re releasing the first set of long-horizon tasks (LHT)<\/b>, which are tasks of great complexity that take an engineer multiple days or even a week to complete. We are also introducing agentic evaluation, starting with agents from corresponding model providers. This addition brings us to <b><a href=\"http:\/\/d.android.com\/bench\">Android Bench 2.0<\/a><\/b>\u2014a major upgrade designed to evaluate AI models and agents against the scale, ambiguity, and complex multi-step problem solving that you tackle every day.<\/p>\n<div class=\"separator\" style=\"clear: both; text-align: center;\"><a href=\"http:\/\/d.android.com\/bench\" style=\"margin-left: 1em; margin-right: 1em;\" target=\"_blank\"><img data-recalc-dims=\"1\" decoding=\"async\" border=\"0\" data-original-height=\"1703\" data-original-width=\"2536\" src=\"https:\/\/i0.wp.com\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEiDNyWS8ZcGbPxKHKh-hH2zRWM5rA2vx6LZB5MMhXiTy3k866YsHg0yc-AhRZlELDH5-qxmSzzNq58ZDEjQ5KRLUY66SaW2XeTQTMuO7eLZ-qf2ue4iygTHglQsABrWESQCgi0BTuvgbMwXMRfiNUTu2bZYfZc6r30U_rPWF6rEktBBTQykmOh7xhz_-zw\/s1600\/LeaderboardFinal%20%281%29.png?w=1280&#038;ssl=1\"><\/a><\/div>\n<\/p>\n<p><\/p>\n<div style=\"text-align: center;\">\n  <i>The Android Bench 2.0 leaderboard<\/i>\n<\/div>\n<h2 style=\"margin-top: 20px;\">From incremental fixes to long-horizon tasks<\/h2>\n<p>The first iteration of Android Bench, along with similar early AI coding benchmarks, focused on incremental changes to existing repositories, in many cases limited to bug fixes or smaller feature requests. This was a reflection of the capabilities of AI assistance at the time, as well as how you were using it. <\/p>\n<p>To continue helping you find the models and coding agents best suited to your development workflow, we have raised the bar of our evaluations to match the work you delegate to AI. Android Bench 2.0 mirrors these ambitious challenges with LHTs that include upgrading dependencies, adding new features, building apps from scratch, or converting a cross-platform app to Android.<\/p>\n<h2>Complex tasks require a more nuanced evaluation and scoring<\/h2>\n<p>On multi-day engineering tasks, binary pass or fail grading doesn\u2019t capture the full picture.<\/p>\n<p>For example, an agent might refactor 40 screens to Jetpack Compose, set up database tables, and pass 90% of requirements, but fail a single edge-case assertion. Binary scoring rates this run as 0%, obscuring the model&#8217;s architectural capabilities. We are moving to continuous scoring to provide a more meaningful signal, both for model development and for your understanding of how AI can help you.<\/p>\n<p>We calculate this completion rate through a combination of factors like functionality, visual fidelity, and avoiding regressions. We also apply objective scoring penalties for deviations from evaluation instructions or structural constraints. Check out the updated leaderboard and click into each model\u2019s card view to see additional elements such as the pass rate, completion rate, and average costs per model and per task. <\/p>\n<p>\n<b>The highest pass rate for LHTs is around 28%<\/b>, much lower than the ~91% for the original tasks in the benchmark. <\/p>\n<div class=\"separator\" style=\"clear: both; text-align: center;\"><a href=\"https:\/\/i0.wp.com\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEj3aYqtv_Zuu08BVYvhfMphyphenhyphen6QBC8V6KLn3Qk6jtfPJdvet3WWU1_rttt1_qraEPpjLLopvdWKVUmHG0VCQ77u12PaGlzinFivVWeABACT5XKdfva0A892EWW5_yd1K_6-fHwvL_7ypGcEWulnnENMKiExu9nOm8jsR59fx3PCejaSWKxF83GYe2LNAYxI\/s1462\/Screenshot%202026-09-16%20at%203.18.23%E2%80%AFPM.png?ssl=1\" style=\"margin-left: 1em; margin-right: 1em;\"><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" border=\"0\" data-original-height=\"1462\" data-original-width=\"1230\" height=\"1256\" src=\"https:\/\/i0.wp.com\/blogger.googleusercontent.com\/img\/b\/R29vZ2xl\/AVvXsEj3aYqtv_Zuu08BVYvhfMphyphenhyphen6QBC8V6KLn3Qk6jtfPJdvet3WWU1_rttt1_qraEPpjLLopvdWKVUmHG0VCQ77u12PaGlzinFivVWeABACT5XKdfva0A892EWW5_yd1K_6-fHwvL_7ypGcEWulnnENMKiExu9nOm8jsR59fx3PCejaSWKxF83GYe2LNAYxI\/w1056-h1256\/Screenshot%202026-09-16%20at%203.18.23%E2%80%AFPM.png?resize=1056%2C1256&#038;ssl=1\" width=\"1056\"><\/a><\/div>\n<div style=\"text-align: center;\"><i>The model card view allows you to explore the strengths and pitfalls of each model<\/i><\/div>\n<div style=\"text-align: center;\"><i><br \/><\/i><\/div>\n<h2 style=\"margin-top: 0px;\">Long-horizon tasks uncover helpful insights for AI assistance<\/h2>\n<p>Beyond measuring how well AI handles long-running tasks, the LHT dataset helps us learn more about the strengths and weaknesses of tested models, and we offer you more practical guidance.<\/p>\n<p>Across model tiers, AI does a better job at writing new code rather than refactoring existing code. Refactors and migrations get trickier because success depends on architectural complexity rather than code volume.<\/p>\n<p>Models show strong capabilities on well-established, deterministic transformations, such as converting Java to Kotlin, swapping Retrofit for Ktor, or introducing a ViewModel layer. They apply these patterns consistently, even across 125+ files and 8,000+ lines of code. <\/p>\n<p>However, models struggle when tasks require runtime validation (like missing dependency injection graphs), involve breaking framework changes, or run into knowledge gaps with unreleased libraries. Porting cross-platform apps to Android remains an open challenge\u2014no model hits a 100% pass rate, and frontier models reach at most a 80% completion rate.<\/p>\n<h2>Introducing agent evaluations<\/h2>\n<p>To help you get a better sense of how models perform when integrated into your agentic workflows, we are adding commonly used agents into our evaluation. We&#8217;re starting by running new models against LHTs with agents from the corresponding model provider. For example, we ran GPT 5.6 Sol on Codex, and Gemini 3.8 Flash on Google Antigravity. This pairing shows how harness design positively impacts developer outcomes, as we\u2019ve seen prompt caching and compact tool windowing can result in token reductions.<\/p>\n<p>We\u2019ll be expanding this in the future by also highlighting results across various model and agent combinations, to help you discover which combinations work best for you and your team. <\/p>\n<p>We invest in this measurement because it\u2019s important for you to be able to use your agent and model of choice for Android development, and we&#8217;ll have more to share with you in the coming weeks.<\/p>\n<h2>New models added<\/h2>\n<p>In addition, we are continuing to expand our leaderboard to ensure you have the most up-to-date data for your development decisions. We added Gemini 3.8 Flash, Gemini 3.7 Flash, OpenAI\u2019s GPT-6, Anthropic\u2019s Fable 5.1, Kimi K3, and Qwen 3.8 Max, with <b>OpenAI\u2019s GPT-6 Astra at the top with a 28% pass rate<\/b>.<\/p>\n<h2>Looking ahead<\/h2>\n<p>Android Bench 2.0 delivers a robust environment for measuring AI for Android development. By combining long-horizon tasks, multimodal evaluation, agents, and continuous scoring, we hope to empower AI research teams to build more capable, dependable AI coding partners, and we hope to provide you with more transparency about your options for AI development. <\/p>\n<p>Check out the <a href=\"http:\/\/d.android.com\/bench\" target=\"_blank\">updated leaderboard<\/a> along with the <a href=\"https:\/\/developer.android.com\/bench\/methodology\/2\">updated methodology<\/a>. Your feedback directly influences how we evolve Android Bench, so please continue to share your feedback with us on <a href=\"https:\/\/github.com\/android-bench\/community-dataset\" target=\"_blank\">GitHub<\/a>, as well as our social channels like <a href=\"https:\/\/x.com\/AndroidDev\" target=\"_blank\">X<\/a> and <a href=\"https:\/\/www.linkedin.com\/showcase\/androiddev\/\" target=\"_blank\">LinkedIn<\/a>.<\/p>\n<\/div>\n<p><\/div>\n<div class=\"pvc_clear\"><\/div>\n<p id=\"pvc_stats_3635\" class=\"pvc_stats total_only\" data-element-id=\"3635\" style=\"\"><i class=\"pvc-stats-icon medium\" aria-hidden=\"true\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" version=\"1.0\" viewbox=\"0 0 502 315\" preserveaspectratio=\"xMidYMid meet\"><g transform=\"translate(0,332) scale(0.1,-0.1)\" fill=\"\" stroke=\"none\"><path d=\"M2394 3279 l-29 -30 -3 -207 c-2 -182 0 -211 15 -242 39 -76 157 -76 196 0 15 31 17 60 15 243 l-3 209 -33 29 c-26 23 -41 29 -80 29 -41 0 -53 -5 -78 -31z\"\/><path d=\"M3085 3251 c-45 -19 -58 -50 -96 -229 -47 -217 -49 -260 -13 -295 52 -53 146 -42 177 20 16 31 87 366 87 410 0 70 -86 122 -155 94z\"\/><path d=\"M1751 3234 c-13 -9 -29 -31 -37 -50 -12 -29 -10 -49 21 -204 19 -94 39 -189 45 -210 14 -50 54 -80 110 -80 34 0 48 6 76 34 21 21 34 44 34 59 0 14 -18 113 -40 219 -37 178 -43 195 -70 221 -36 32 -101 37 -139 11z\"\/><path d=\"M1163 3073 c-36 -7 -73 -59 -73 -102 0 -56 133 -378 171 -413 34 -32 83 -37 129 -13 70 36 67 87 -16 290 -86 209 -89 214 -129 231 -35 14 -42 15 -82 7z\"\/><path d=\"M3689 3066 c-15 -9 -33 -30 -42 -48 -48 -103 -147 -355 -147 -375 0 -98 131 -148 192 -74 13 15 57 108 97 206 80 196 84 226 37 273 -30 30 -99 39 -137 18z\"\/><path d=\"M583 2784 c-38 -19 -67 -74 -58 -113 9 -42 211 -354 242 -373 16 -10 45 -18 66 -18 51 0 107 52 107 100 0 39 -1 41 -124 234 -80 126 -108 162 -133 173 -41 17 -61 16 -100 -3z\"\/><path d=\"M4250 2784 c-14 -9 -74 -91 -133 -183 -95 -150 -107 -173 -107 -213 0 -55 33 -94 87 -104 67 -13 90 8 211 198 130 202 137 225 78 284 -27 27 -42 34 -72 34 -22 0 -50 -8 -64 -16z\"\/><path d=\"M2275 2693 c-553 -48 -1095 -270 -1585 -649 -135 -104 -459 -423 -483 -476 -23 -49 -22 -139 2 -186 73 -142 361 -457 571 -626 285 -228 642 -407 990 -497 242 -63 336 -73 660 -74 310 0 370 5 595 52 535 111 1045 392 1455 803 122 121 250 273 275 326 19 41 19 137 0 174 -41 79 -309 363 -465 492 -447 370 -946 591 -1479 653 -113 14 -422 18 -536 8z m395 -428 c171 -34 330 -124 456 -258 112 -119 167 -219 211 -378 27 -96 24 -300 -5 -401 -72 -255 -236 -447 -474 -557 -132 -62 -201 -76 -368 -76 -167 0 -236 14 -368 76 -213 98 -373 271 -451 485 -162 444 86 934 547 1084 153 49 292 57 452 25z m909 -232 c222 -123 408 -262 593 -441 76 -74 138 -139 138 -144 0 -16 -233 -242 -330 -319 -155 -123 -309 -223 -461 -299 l-81 -41 32 46 c18 26 49 83 70 128 143 306 141 649 -6 957 -25 52 -61 116 -79 142 l-34 47 45 -20 c26 -10 76 -36 113 -56z m-2057 25 c-40 -58 -105 -190 -130 -263 -110 -324 -59 -707 132 -981 25 -35 42 -64 37 -64 -19 0 -241 119 -326 174 -188 122 -406 314 -532 468 l-58 71 108 103 c185 178 428 349 672 473 66 33 121 60 123 61 2 0 -10 -19 -26 -42z\"\/><path d=\"M2375 1950 c-198 -44 -350 -190 -395 -379 -18 -76 -8 -221 19 -290 114 -284 457 -406 731 -260 98 52 188 154 231 260 27 69 37 214 19 290 -38 163 -166 304 -326 360 -67 23 -215 33 -279 19z\"\/><\/g><\/svg><\/i> <img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" width=\"16\" height=\"16\" alt=\"Loading\" src=\"https:\/\/i0.wp.com\/dev95.site\/wp-content\/plugins\/page-views-count\/ajax-loader-2x.gif?resize=16%2C16&#038;ssl=1\" border=\"0\" \/><\/p>\n<div class=\"pvc_clear\"><\/div>","protected":false},"excerpt":{"rendered":"<p>Posted by Matthew McCullough, VP, Product Management, Android Developer When we first launched Android Bench, we built a rigorous foundation for evaluating how large language models (LLMs) assist developers with real-world Android tasks. As AI models and agents rapidly evolve,<\/p>\n<div class=\"hosteria-entry-more\"><a href=\"https:\/\/dev95.site\/ar\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/\" class=\"no-underline font-light  group-hover:text-primary-800 dark:group-hover:text-primary-300 py-1\">Read more &gt;&gt;&gt;<\/a><\/div>\n<div class=\"pvc_clear\"><\/div>\n<p id=\"pvc_stats_3635\" class=\"pvc_stats total_only\" data-element-id=\"3635\" style=\"\"><i class=\"pvc-stats-icon medium\" aria-hidden=\"true\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" version=\"1.0\" viewbox=\"0 0 502 315\" preserveaspectratio=\"xMidYMid meet\"><g transform=\"translate(0,332) scale(0.1,-0.1)\" fill=\"\" stroke=\"none\"><path d=\"M2394 3279 l-29 -30 -3 -207 c-2 -182 0 -211 15 -242 39 -76 157 -76 196 0 15 31 17 60 15 243 l-3 209 -33 29 c-26 23 -41 29 -80 29 -41 0 -53 -5 -78 -31z\"\/><path d=\"M3085 3251 c-45 -19 -58 -50 -96 -229 -47 -217 -49 -260 -13 -295 52 -53 146 -42 177 20 16 31 87 366 87 410 0 70 -86 122 -155 94z\"\/><path d=\"M1751 3234 c-13 -9 -29 -31 -37 -50 -12 -29 -10 -49 21 -204 19 -94 39 -189 45 -210 14 -50 54 -80 110 -80 34 0 48 6 76 34 21 21 34 44 34 59 0 14 -18 113 -40 219 -37 178 -43 195 -70 221 -36 32 -101 37 -139 11z\"\/><path d=\"M1163 3073 c-36 -7 -73 -59 -73 -102 0 -56 133 -378 171 -413 34 -32 83 -37 129 -13 70 36 67 87 -16 290 -86 209 -89 214 -129 231 -35 14 -42 15 -82 7z\"\/><path d=\"M3689 3066 c-15 -9 -33 -30 -42 -48 -48 -103 -147 -355 -147 -375 0 -98 131 -148 192 -74 13 15 57 108 97 206 80 196 84 226 37 273 -30 30 -99 39 -137 18z\"\/><path d=\"M583 2784 c-38 -19 -67 -74 -58 -113 9 -42 211 -354 242 -373 16 -10 45 -18 66 -18 51 0 107 52 107 100 0 39 -1 41 -124 234 -80 126 -108 162 -133 173 -41 17 -61 16 -100 -3z\"\/><path d=\"M4250 2784 c-14 -9 -74 -91 -133 -183 -95 -150 -107 -173 -107 -213 0 -55 33 -94 87 -104 67 -13 90 8 211 198 130 202 137 225 78 284 -27 27 -42 34 -72 34 -22 0 -50 -8 -64 -16z\"\/><path d=\"M2275 2693 c-553 -48 -1095 -270 -1585 -649 -135 -104 -459 -423 -483 -476 -23 -49 -22 -139 2 -186 73 -142 361 -457 571 -626 285 -228 642 -407 990 -497 242 -63 336 -73 660 -74 310 0 370 5 595 52 535 111 1045 392 1455 803 122 121 250 273 275 326 19 41 19 137 0 174 -41 79 -309 363 -465 492 -447 370 -946 591 -1479 653 -113 14 -422 18 -536 8z m395 -428 c171 -34 330 -124 456 -258 112 -119 167 -219 211 -378 27 -96 24 -300 -5 -401 -72 -255 -236 -447 -474 -557 -132 -62 -201 -76 -368 -76 -167 0 -236 14 -368 76 -213 98 -373 271 -451 485 -162 444 86 934 547 1084 153 49 292 57 452 25z m909 -232 c222 -123 408 -262 593 -441 76 -74 138 -139 138 -144 0 -16 -233 -242 -330 -319 -155 -123 -309 -223 -461 -299 l-81 -41 32 46 c18 26 49 83 70 128 143 306 141 649 -6 957 -25 52 -61 116 -79 142 l-34 47 45 -20 c26 -10 76 -36 113 -56z m-2057 25 c-40 -58 -105 -190 -130 -263 -110 -324 -59 -707 132 -981 25 -35 42 -64 37 -64 -19 0 -241 119 -326 174 -188 122 -406 314 -532 468 l-58 71 108 103 c185 178 428 349 672 473 66 33 121 60 123 61 2 0 -10 -19 -26 -42z\"\/><path d=\"M2375 1950 c-198 -44 -350 -190 -395 -379 -18 -76 -8 -221 19 -290 114 -284 457 -406 731 -260 98 52 188 154 231 260 27 69 37 214 19 290 -38 163 -166 304 -326 360 -67 23 -215 33 -279 19z\"\/><\/g><\/svg><\/i> <img loading=\"lazy\" decoding=\"async\" width=\"16\" height=\"16\" alt=\"Loading\" src=\"https:\/\/dev95.site\/wp-content\/plugins\/page-views-count\/ajax-loader-2x.gif\" border=\"0\" \/><\/p>\n<div class=\"pvc_clear\"><\/div>","protected":false},"author":1,"featured_media":3636,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"fp_fajr_begins":"","fp_fajr_iqamah":"","fp_dhuhr_begins":"","fp_dhuhr_iqamah":"","fp_asr_begins":"","fp_asr_iqamah":"","fp_maghrib_begins":"","fp_maghrib_iqamah":"","fp_isha_begins":"","fp_isha_iqamah":"","fp_midnight":"","fp_midnight_name":"","fp_sunrise":"","fp_single_prayer_begins_title":"","fp_single_prayer_iqamah_title":"","fp_prayer_times_for_today":"","fp_hijra_date":"","fp_fajr_name":"","fp_dhuhr_name":"","fp_asr_name":"","fp_maghrib_name":"","fp_isha_name":"","fp_sunrise_name":"","fp_currentDate":"","fp_current_time":"","fp_current_title":"","fp_current_location":"","fp_masjid_name":"","fp_prayer_title":"","fp_next_prayer_iqamah_time":"","fp_next_prayer_iqamah_title":"","fp_next_prayer_begins_time":"","fp_next_prayer_begins_title":"","fp_next_prayer_title":"","_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_post_was_ever_published":false},"categories":[9],"tags":[],"class_list":["post-3635","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-android"],"a3_pvc":{"activated":true,"total_views":0,"today_views":0},"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.5 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Android Bench 2.0: Pushing the frontier with challenging long-horizon tasks - Dev95<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/dev95.site\/ar\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/\" \/>\n<meta property=\"og:locale\" content=\"ar_AR\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Android Bench 2.0: Pushing the frontier with challenging long-horizon tasks - Dev95\" \/>\n<meta property=\"og:description\" content=\"Posted by Matthew McCullough, VP, Product Management, Android Developer When we first launched Android Bench, we built a rigorous foundation for evaluating how large language models (LLMs) assist developers with real-world Android tasks. As AI models and agents rapidly evolve,Read more &gt;&gt;&gt;\" \/>\n<meta property=\"og:url\" content=\"https:\/\/dev95.site\/ar\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/\" \/>\n<meta property=\"og:site_name\" content=\"Dev95\" \/>\n<meta property=\"article:published_time\" content=\"2026-09-23T10:14:27+00:00\" \/>\n<meta name=\"author\" content=\"dev95\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"\u0643\u064f\u062a\u0628 \u0628\u0648\u0627\u0633\u0637\u0629\" \/>\n\t<meta name=\"twitter:data1\" content=\"dev95\" \/>\n\t<meta name=\"twitter:label2\" content=\"\u0648\u0642\u062a \u0627\u0644\u0642\u0631\u0627\u0621\u0629 \u0627\u0644\u0645\u064f\u0642\u062f\u0651\u0631\" \/>\n\t<meta name=\"twitter:data2\" content=\"4 \u062f\u0642\u0627\u0626\u0642\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/dev95.site\\\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/dev95.site\\\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\\\/\"},\"author\":{\"name\":\"dev95\",\"@id\":\"https:\\\/\\\/dev95.site\\\/#\\\/schema\\\/person\\\/b807805ffe2916206b04d0938bce0298\"},\"headline\":\"Android Bench 2.0: Pushing the frontier with challenging long-horizon tasks\",\"datePublished\":\"2026-09-23T10:14:27+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/dev95.site\\\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\\\/\"},\"wordCount\":895,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\\\/\\\/dev95.site\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/dev95.site\\\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/i0.wp.com\\\/dev95.site\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/Bench20Metadata-bench.png?fit=2049%2C1324&ssl=1\",\"articleSection\":[\"Android\"],\"inLanguage\":\"ar\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\\\/\\\/dev95.site\\\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\\\/#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/dev95.site\\\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\\\/\",\"url\":\"https:\\\/\\\/dev95.site\\\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\\\/\",\"name\":\"Android Bench 2.0: Pushing the frontier with challenging long-horizon tasks - Dev95\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/dev95.site\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/dev95.site\\\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/dev95.site\\\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/i0.wp.com\\\/dev95.site\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/Bench20Metadata-bench.png?fit=2049%2C1324&ssl=1\",\"datePublished\":\"2026-09-23T10:14:27+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/dev95.site\\\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\\\/#breadcrumb\"},\"inLanguage\":\"ar\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/dev95.site\\\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"ar\",\"@id\":\"https:\\\/\\\/dev95.site\\\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\\\/#primaryimage\",\"url\":\"https:\\\/\\\/i0.wp.com\\\/dev95.site\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/Bench20Metadata-bench.png?fit=2049%2C1324&ssl=1\",\"contentUrl\":\"https:\\\/\\\/i0.wp.com\\\/dev95.site\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/Bench20Metadata-bench.png?fit=2049%2C1324&ssl=1\",\"width\":2049,\"height\":1324},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/dev95.site\\\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/dev95.site\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Android Bench 2.0: Pushing the frontier with challenging long-horizon tasks\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/dev95.site\\\/#website\",\"url\":\"https:\\\/\\\/dev95.site\\\/\",\"name\":\"Dev95\",\"description\":\"A comprehensive platform for data and knowledge, delivering reliable content that meets the aspirations of readers and enthusiasts.\",\"publisher\":{\"@id\":\"https:\\\/\\\/dev95.site\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/dev95.site\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"ar\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/dev95.site\\\/#organization\",\"name\":\"Dev95\",\"url\":\"https:\\\/\\\/dev95.site\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"ar\",\"@id\":\"https:\\\/\\\/dev95.site\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/i0.wp.com\\\/dev95.site\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/rbrrbr-6.png?fit=512%2C512&ssl=1\",\"contentUrl\":\"https:\\\/\\\/i0.wp.com\\\/dev95.site\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/rbrrbr-6.png?fit=512%2C512&ssl=1\",\"width\":512,\"height\":512,\"caption\":\"Dev95\"},\"image\":{\"@id\":\"https:\\\/\\\/dev95.site\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/dev95.site\\\/#\\\/schema\\\/person\\\/b807805ffe2916206b04d0938bce0298\",\"name\":\"dev95\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"ar\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/a70a73d950838b20cd80d7ebdc955737e802e8cd896044c5473b32b946c0662a?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/a70a73d950838b20cd80d7ebdc955737e802e8cd896044c5473b32b946c0662a?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/a70a73d950838b20cd80d7ebdc955737e802e8cd896044c5473b32b946c0662a?s=96&d=mm&r=g\",\"caption\":\"dev95\"},\"url\":\"https:\\\/\\\/dev95.site\\\/ar\\\/author\\\/mohammad\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Android Bench 2.0: Pushing the frontier with challenging long-horizon tasks - Dev95","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/dev95.site\/ar\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/","og_locale":"ar_AR","og_type":"article","og_title":"Android Bench 2.0: Pushing the frontier with challenging long-horizon tasks - Dev95","og_description":"Posted by Matthew McCullough, VP, Product Management, Android Developer When we first launched Android Bench, we built a rigorous foundation for evaluating how large language models (LLMs) assist developers with real-world Android tasks. As AI models and agents rapidly evolve,Read more &gt;&gt;&gt;","og_url":"https:\/\/dev95.site\/ar\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/","og_site_name":"Dev95","article_published_time":"2026-09-23T10:14:27+00:00","author":"dev95","twitter_card":"summary_large_image","twitter_misc":{"\u0643\u064f\u062a\u0628 \u0628\u0648\u0627\u0633\u0637\u0629":"dev95","\u0648\u0642\u062a \u0627\u0644\u0642\u0631\u0627\u0621\u0629 \u0627\u0644\u0645\u064f\u0642\u062f\u0651\u0631":"4 \u062f\u0642\u0627\u0626\u0642"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/dev95.site\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/#article","isPartOf":{"@id":"https:\/\/dev95.site\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/"},"author":{"name":"dev95","@id":"https:\/\/dev95.site\/#\/schema\/person\/b807805ffe2916206b04d0938bce0298"},"headline":"Android Bench 2.0: Pushing the frontier with challenging long-horizon tasks","datePublished":"2026-09-23T10:14:27+00:00","mainEntityOfPage":{"@id":"https:\/\/dev95.site\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/"},"wordCount":895,"commentCount":0,"publisher":{"@id":"https:\/\/dev95.site\/#organization"},"image":{"@id":"https:\/\/dev95.site\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/#primaryimage"},"thumbnailUrl":"https:\/\/i0.wp.com\/dev95.site\/wp-content\/uploads\/2026\/09\/Bench20Metadata-bench.png?fit=2049%2C1324&ssl=1","articleSection":["Android"],"inLanguage":"ar","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/dev95.site\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/dev95.site\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/","url":"https:\/\/dev95.site\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/","name":"Android Bench 2.0: Pushing the frontier with challenging long-horizon tasks - Dev95","isPartOf":{"@id":"https:\/\/dev95.site\/#website"},"primaryImageOfPage":{"@id":"https:\/\/dev95.site\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/#primaryimage"},"image":{"@id":"https:\/\/dev95.site\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/#primaryimage"},"thumbnailUrl":"https:\/\/i0.wp.com\/dev95.site\/wp-content\/uploads\/2026\/09\/Bench20Metadata-bench.png?fit=2049%2C1324&ssl=1","datePublished":"2026-09-23T10:14:27+00:00","breadcrumb":{"@id":"https:\/\/dev95.site\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/#breadcrumb"},"inLanguage":"ar","potentialAction":[{"@type":"ReadAction","target":["https:\/\/dev95.site\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/"]}]},{"@type":"ImageObject","inLanguage":"ar","@id":"https:\/\/dev95.site\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/#primaryimage","url":"https:\/\/i0.wp.com\/dev95.site\/wp-content\/uploads\/2026\/09\/Bench20Metadata-bench.png?fit=2049%2C1324&ssl=1","contentUrl":"https:\/\/i0.wp.com\/dev95.site\/wp-content\/uploads\/2026\/09\/Bench20Metadata-bench.png?fit=2049%2C1324&ssl=1","width":2049,"height":1324},{"@type":"BreadcrumbList","@id":"https:\/\/dev95.site\/android-bench-2-0-pushing-the-frontier-with-challenging-long-horizon-tasks\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/dev95.site\/"},{"@type":"ListItem","position":2,"name":"Android Bench 2.0: Pushing the frontier with challenging long-horizon tasks"}]},{"@type":"WebSite","@id":"https:\/\/dev95.site\/#website","url":"https:\/\/dev95.site\/","name":"Dev95","description":"A comprehensive platform for data and knowledge, delivering reliable content that meets the aspirations of readers and enthusiasts.","publisher":{"@id":"https:\/\/dev95.site\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/dev95.site\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"ar"},{"@type":"Organization","@id":"https:\/\/dev95.site\/#organization","name":"Dev95","url":"https:\/\/dev95.site\/","logo":{"@type":"ImageObject","inLanguage":"ar","@id":"https:\/\/dev95.site\/#\/schema\/logo\/image\/","url":"https:\/\/i0.wp.com\/dev95.site\/wp-content\/uploads\/2026\/07\/rbrrbr-6.png?fit=512%2C512&ssl=1","contentUrl":"https:\/\/i0.wp.com\/dev95.site\/wp-content\/uploads\/2026\/07\/rbrrbr-6.png?fit=512%2C512&ssl=1","width":512,"height":512,"caption":"Dev95"},"image":{"@id":"https:\/\/dev95.site\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/dev95.site\/#\/schema\/person\/b807805ffe2916206b04d0938bce0298","name":"dev95","image":{"@type":"ImageObject","inLanguage":"ar","@id":"https:\/\/secure.gravatar.com\/avatar\/a70a73d950838b20cd80d7ebdc955737e802e8cd896044c5473b32b946c0662a?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/a70a73d950838b20cd80d7ebdc955737e802e8cd896044c5473b32b946c0662a?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/a70a73d950838b20cd80d7ebdc955737e802e8cd896044c5473b32b946c0662a?s=96&d=mm&r=g","caption":"dev95"},"url":"https:\/\/dev95.site\/ar\/author\/mohammad\/"}]}},"jetpack_sharing_enabled":true,"jetpack_featured_media_url":"https:\/\/i0.wp.com\/dev95.site\/wp-content\/uploads\/2026\/09\/Bench20Metadata-bench.png?fit=2049%2C1324&ssl=1","_links":{"self":[{"href":"https:\/\/dev95.site\/ar\/wp-json\/wp\/v2\/posts\/3635","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/dev95.site\/ar\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/dev95.site\/ar\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/dev95.site\/ar\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/dev95.site\/ar\/wp-json\/wp\/v2\/comments?post=3635"}],"version-history":[{"count":0,"href":"https:\/\/dev95.site\/ar\/wp-json\/wp\/v2\/posts\/3635\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/dev95.site\/ar\/wp-json\/wp\/v2\/media\/3636"}],"wp:attachment":[{"href":"https:\/\/dev95.site\/ar\/wp-json\/wp\/v2\/media?parent=3635"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/dev95.site\/ar\/wp-json\/wp\/v2\/categories?post=3635"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/dev95.site\/ar\/wp-json\/wp\/v2\/tags?post=3635"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}