Close Menu
TechCentralTechCentral

    Subscribe to the newsletter

    Get the best South African technology news and analysis delivered to your e-mail inbox every morning.

    Facebook X (Twitter) YouTube LinkedIn
    WhatsApp Facebook X (Twitter) LinkedIn YouTube
    TechCentralTechCentral
    • News
      Nedbank hires MTN's former tech chief as group CIO - Nikos Angelopoulos

      Nedbank hires MTN’s former tech chief as group CIO

      31 July 2026
      Eskom's diesel bill falls 86% as breakdowns hit eight-year low

      Eskom’s diesel bill falls 86% as breakdowns hit eight-year low

      31 July 2026
      Ramaphosa signs off on taking the grid away from Eskom

      Ramaphosa signs off on taking the grid away from Eskom

      31 July 2026
      Microsoft just had the biggest day in stock market history

      Microsoft just had the biggest day in stock market history

      31 July 2026
      MTN Nigeria's growth engine stalled in second quarter - Karl Toriola

      MTN Nigeria’s growth engine stalled in second quarter

      31 July 2026
    • World
      Meta AI will now tell parents if their teen is in crisis

      Meta AI will now tell parents if their teen is in crisis

      17 July 2026
      IBM shares crash 25% as AI upends software spending - Arvind Krishna

      IBM shares crash 25% as AI upends software spending

      15 July 2026
      Jony Ive's first OpenAI device: an AI smart speaker - Jony Ive and Sam Altman

      Jony Ive’s first OpenAI device: an AI smart speaker

      15 July 2026
      Stripe, Advent in talks to buy PayPal for $53-billion

      Stripe, Advent in talks to buy PayPal for $53-billion

      15 July 2026
      Memory crisis sends smartphone market into steep decline

      Memory crisis sends smartphone market into steep decline

      13 July 2026
    • In-depth
      The plan to stop AI from breaking the world - Google DeepMind CEO Demis Hassabis. Image: John Sears

      The plan to stop AI from breaking the world

      16 July 2026
      The internet has a Strait of Hormuz problem

      The internet has a Strait of Hormuz problem

      15 July 2026
      AI boom sparks rally, frenzy and fear

      AI boom sparks rally, frenzy and fear

      11 June 2026
      Every plug-in hybrid on sale in South Africa, ranked by price - Lamborghini Temerario

      Every plug-in hybrid on sale in South Africa, ranked by price

      7 June 2026
      What Wi-Fi 8 will mean for wireless networks

      What Wi-Fi 8 will mean for wireless networks

      1 June 2026
    • TCS
      TCS+ | Why South African workers must become supervisors of digital labour - Accelera Digital Group Cliff de Wit

      TCS+ | Why South African workers must become supervisors of digital labour

      31 July 2026
      TCS | Rapid deployment rules can't work without municipalities: ACT - Nomvuyiso Batyi

      TCS | Icasa’s rules skip the real bottleneck: ACT

      30 July 2026
      TCS+ | iStore Business on why Apple makes sense for SMEs - Sudesh Pillay and Tamia Nontsikelelo

      TCS+ | iStore Business on why Apple makes sense for SMEs

      30 July 2026
      TCS+ | A smarter approach to cloud for South African businesses - Joel Chacko and Jonathan Oaker

      TCS+ | A smarter approach to cloud for South African businesses

      28 July 2026
      TCS | How Optasia lends billions to people banks can't see - Salvador Anglada

      TCS | How Optasia lends billions to people banks can’t see

      23 July 2026
    • Opinion
      The author, Jannie van Zyl

      Selling vapour is corporate suicide in slow motion

      16 July 2026
      Brazil's online gambling crackdown is a lesson for South Africa

      How Amazon outmanoeuvred Starlink in South Africa

      15 July 2026
      The Popia problem with agentic AI - Herman Haasbroek

      The Popia problem with agentic AI

      14 July 2026
      The author, Fanie van Rooyen

      South Africa can still catch the AI wave – here’s how

      7 July 2026
      The author, Fanie van Rooyen

      The AI utopia South Africa can’t afford

      1 July 2026
    • Company Hubs
      • 1Stream
      • Africa Data Centres
      • AfriGIS
      • Altron Digital Business
      • Altron Document Solutions
      • Altron Group
      • Arctic Wolf
      • Ascent Technology
      • AvertITD
      • BBD
      • Braintree
      • CallMiner
      • CambriLearn
      • CM Telecom
      • Contactable
      • CYBER1 Solutions
      • Digicloud Africa
      • Digimune
      • Domains.co.za
      • ESET
      • Euphoria Telecom
      • HOSTAFRICA
      • Incredible Business
      • iONLINE
      • IQbusiness
      • Iris Network Systems
      • Kaspersky
      • LSD Open
      • Mitel
      • NEC XON
      • Netstar
      • Network Platforms
      • Next DLP
      • Ovations
      • Paracon
      • Paratus
      • Q-KON
      • SevenC
      • SkyWire
      • Solid8 Technologies
      • Telit Cinterion
      • Telviva
      • Tenable
      • Vertiv
      • Videri Digital
      • Vodacom Business
      • Vox
      • Wipro
      • Workday
      • XLink
    • Sections
      • AI and machine learning
      • Banking
      • Broadcasting and Media
      • Cloud services
      • Contact centres and CX
      • Cryptocurrencies
      • Education and skills
      • Electronics and hardware
      • Energy and sustainability
      • Enterprise software
      • Financial services
      • HealthTech
      • Information security
      • Internet and connectivity
      • Internet of Things
      • Investment
      • IT services
      • Lifestyle
      • Policy and regulation
      • Public sector
      • Retail and e-commerce
      • Satellite communications
      • Science
      • SMEs and start-ups
      • Social media
      • Talent and leadership
      • Telecoms
      • Watts & Wheels
    • Events
    • Advertise
    TechCentralTechCentral
    Home » Sections » AI and machine learning » AI Barometer: the best AI for every job right now

    AI Barometer: the best AI for every job right now

    Our new evidence-based barometer ranks the leading AI tools for work, code, writing, images and video.
    By Fanie van Rooyen13 July 2026
    Twitter LinkedIn Facebook WhatsApp Email Telegram Copy Link
    News Alerts
    WhatsApp

    AI Barometer: the best AI for every job right now

    Barely a month goes by without an AI laboratory declaring its latest model the smartest ever built. For anyone simply trying to keep up, or trying to pick a tool that drafts a decent e-mail, fixes broken code or turns a prompt into a usable image, the noise is deafening.

    This is the first instalment of TechCentral’s AI Barometer, a recurring, evidence-based snapshot of which AI models and platforms lead in the tasks people actually use them for. No vibes and no vendor marketing: every verdict below is drawn from public, independently run benchmarks and leaderboards, each chosen because it tests the use case in question.

    The scores were checked in the first week of July 2026 and, given the pace of releases, some will likely have shifted by the time you read this. That churn is precisely why the barometer exists.

    How the barometer works

    Two kinds of evidence feed this report. The first is arena-style testing: platforms such as Arena (until January known as LMArena, an offshoot of research at the University of California, Berkeley) and Artificial Analysis show users two anonymised outputs for the same prompt and ask which is better. Millions of blind votes are converted into chess-style Elo ratings.

    The second is task-based benchmarking, in which models are scored against fixed sets of difficult, realistic problems – resolving real software bugs, answering expert-level science questions or producing the documents and spreadsheets professionals get paid to make.

    Both approaches have known weaknesses. Arena votes reward outputs people like, which is not always the same as outputs that are correct. Task benchmarks can be gamed if test questions leak into a model’s training data, which is why the barometer favours contamination-resistant tests. And when leading models sit within a few points of one another, the honest reading is a statistical tie. Where that applies below, we say so.

    AI

    Everyday assistant and chat

    The broadest measure available is Arena’s text leaderboard, which in the first week of July had collected 7.15-million blind votes across 369 models. Anthropic’s newly released Claude Fable 5 tops the board on an Elo of 1 509 – and, remarkably, Anthropic models fill all five top slots.

    Google’s Gemini 3.1 Pro Preview is the best of the rest on 1 486, with OpenAI’s GPT-5.5 close behind on 1 481.

    The caveat: the top 10 sits within roughly 30 Elo points, so for casual use any of these will serve well. Price separates them more than quality does: on API rates listed by Arena, Gemini 3.1 Pro costs US$2/$12 per million input/output tokens (about R33/R195) against Claude Fable 5’s $10/$50 (about R163/R814), making Gemini the clear value pick.

    That comparison applies to businesses building on the models’ APIs, though; for consumers, the flagship subscriptions from OpenAI, Google and Anthropic are all priced within a few dollars of one another, whatever the underlying token costs.

    Co-working and office tasks

    For AI that does real work – producing the documents, presentations, spreadsheets and analyses of everyday office life – the most relevant public test is GDPval, a dataset of 220 tasks that OpenAI built with industry professionals across 44 occupations. Artificial Analysis runs models through it agentically, with web and shell access, and ranks the deliverables by blind pairwise comparison. Anthropic’s Claude Opus 4.8 leads decisively on an Elo of 1 890, some 121 points clear of GPT-5.5 on 1 769 – a notable result, given that OpenAI created the benchmark.

    Coding

    The barometer’s coding yardstick is SWE-bench Pro, run as an independent public leaderboard by Scale AI. It asks models to resolve real issues from actively maintained software repositories and was designed specifically to resist training-data contamination, a documented problem with the older SWE-bench Verified test. On the independent board, OpenAI’s GPT-5.4 leads on 59.1%, in a statistical tie with Meta’s newly listed Muse Spark on 55%, with Anthropic’s Claude Opus 4.6 next on 51.9%. One flag for readers weighing vendor claims: Anthropic says its new Claude Fable 5 scored 80.3% on the same task set, but that figure was produced with the company’s own tooling and does not yet appear on the independent leaderboard. Until it does, the verified crown belongs to OpenAI.

    Writing

    Writing quality is notoriously hard to score, and the most serious public attempt is the independent EQ-Bench Creative Writing benchmark (version 3), which combines rubric marking with Elo-style pairwise judging across 32 varied prompts. Claude Fable 5 leads on an Elo of 2 192, narrowly ahead of Claude Opus 4.7 on 2 179, with GPT-5.5 third on 2 019. An important caveat applies here: the judge is itself an AI model, so results are directional rather than definitive – though the benchmark also publishes the raw outputs, so sceptical readers can judge the prose for themselves.

    Research

    Research splits into two skills. For deep reasoning, the reference test is Humanity’s Last Exam, 2 500 questions crowdsourced from subject-matter experts precisely because frontier models could not answer them. On the official leaderboard run by Scale AI, Google’s Gemini 3.1 Pro Preview leads on 46.4%, in a statistical tie with OpenAI’s GPT-5.4 Pro on 44.3% – meaning the best models still fail more than half the paper. (Anthropic reports a higher 53.3% for Claude Fable 5, but that score is vendor-reported and not yet on the official board.)

    For live web research, the benchmark to watch is BrowseComp, OpenAI’s 1 266-question test of an agent’s ability to hunt down hard-to-find information online. One honesty note: unlike every other test in this report, BrowseComp has no independently run public leaderboard – the laboratories report their own scores. On those self-reported numbers, OpenAI’s GPT-5.5 Pro is top on 90.1%. In short: Google for depth; ChatGPT, on its own numbers, for tracking things down online.

    When the AI costs more than the coder

    Image generation

    Image models are ranked almost entirely by blind human preference, and the Artificial Analysis Image Arena is the busiest public board. OpenAI’s GPT Image 2 (high) leads on an Elo of 1 339 across more than 13 000 comparisons – 58 points clear of second-placed Reve 2.0 (1 281), with Microsoft’s MAI-Image-2.5 third on 1 271.

    Video generation

    Video is the category where conventional wisdom is most out of date. On the Artificial Analysis Video Arena’s text-to-video board (with audio), ByteDance’s Dreamina Seedance 2.0 leads on 1 224, followed by Alibaba’s Wan 2.7 (1 160) and HappyHorse 1.1 (1 154). Google’s Veo 3.1, the strongest Western entrant, sits only 10th on 1 095.

    On the silent board, Alibaba’s HappyHorse 1.0 leads on 1 290. The top four slots on the board all belong to Chinese developers, and the costs remain steep: producing a minute of 1080p footage runs from about $9 to $24 (roughly R145 to R390) on the model creators’ own APIs.

    Music generation

    Music benchmarking is younger and thinner than any category above, so treat these results as directional. On the Artificial Analysis Music Arena, which uses the same blind-vote method as its image and video boards, Suno v5.5 tops both leaderboards – instrumental (Elo 1 194) and vocals (1 168) – ahead of Mureka V8 in both, with Suno’s earlier V5 holding third place on each board.

    The scoreboard

    Use case Front-runner Runner-up Benchmark used
    Everyday assistant Claude Fable 5 Gemini 3.1 Pro Preview Arena text leaderboard (Elo)
    Co-working/office tasks Claude Opus 4.8 GPT-5.5 GDPval-AA (Elo)
    Coding GPT-5.4* Muse Spark (Meta)* SWE-bench Pro (% resolved)
    Writing Claude Fable 5 Claude Opus 4.7 EQ-Bench Creative Writing v3
    Deep research Gemini 3.1 Pro Preview* GPT-5.4 Pro* Humanity’s Last Exam (%)
    Web research GPT-5.5 Pro† GPT-5.4 Pro† BrowseComp (%, lab-reported)
    Image generation GPT Image 2 Reve 2.0 AA Image Arena (Elo)
    Video generation Seedance 2.0 (ByteDance) Wan 2.7 (Alibaba) AA Video Arena, with audio (Elo)
    Music generation Suno v5.5 Mureka V8 AA Music Arena (Elo)

    * Statistical tie. † Lab-reported; BrowseComp has no independently run public leaderboard. AA = Artificial Analysis. All figures checked in the first week of July 2026

    Until next time

    A barometer measures pressure at a moment in time, and the pressure in this industry rarely holds steady for a quarter. Several of the results above are statistical ties that a month of new releases could flip, and at least two categories – video and music – are moving faster than the benchmarks that measure them. We’ll check back with another reading in two or three months to see what’s changed.  – © 2026 NewsCentral Media

    • Subscribe to TechCentral’s daily newsletter
    • Get breaking news alerts on WhatsApp
    Follow TechCentral on Google News Add TechCentral as your preferred source on Google


    AI Barometer Alibaba Anthropic Artificial Analysis ByteDance ChatGPT Claude Gemini Google GPT-5 Meta Microsoft OpenAI Suno TechCentral AI Barometer
    WhatsApp YouTube
    Share. Facebook Twitter LinkedIn WhatsApp Telegram Email Copy Link
    Previous ArticleShoprite claims early win in grocery AI wars
    Next Article Industry to Icasa: punish municipalities that stall network roll-out

    Related Posts

    Microsoft just had the biggest day in stock market history

    Microsoft just had the biggest day in stock market history

    31 July 2026
    Meta cash flow collapses as AI bill hits $145-billion - Mark Zuckerberg

    Meta cash flow collapses as AI bill hits $145-billion

    30 July 2026
    Africa's most popular phones have a tracking problem

    Africa’s most popular phones have a tracking problem

    24 July 2026
    Company News
    Domains.co.za launches self-hosted n8n VPS hosting

    Domains.co.za launches self-hosted n8n VPS hosting

    31 July 2026
    Smarter.tech '26 shows why smarter technology begins with context - Obsidian Systems

    Context is the missing piece in enterprise AI: Obsidian

    31 July 2026
    Huawei launches 12 intelligent transport solutions in South Africa - Sam Tang

    Huawei launches 12 intelligent transport solutions in South Africa

    30 July 2026
    Opinion
    The author, Jannie van Zyl

    Selling vapour is corporate suicide in slow motion

    16 July 2026
    Brazil's online gambling crackdown is a lesson for South Africa

    How Amazon outmanoeuvred Starlink in South Africa

    15 July 2026
    The Popia problem with agentic AI - Herman Haasbroek

    The Popia problem with agentic AI

    14 July 2026

    Subscribe to Updates

    Get the best South African technology news and analysis delivered to your e-mail inbox every morning.

    Latest Posts
    Nedbank hires MTN's former tech chief as group CIO - Nikos Angelopoulos

    Nedbank hires MTN’s former tech chief as group CIO

    31 July 2026
    Eskom's diesel bill falls 86% as breakdowns hit eight-year low

    Eskom’s diesel bill falls 86% as breakdowns hit eight-year low

    31 July 2026
    Ramaphosa signs off on taking the grid away from Eskom

    Ramaphosa signs off on taking the grid away from Eskom

    31 July 2026
    TCS+ | Why South African workers must become supervisors of digital labour - Accelera Digital Group Cliff de Wit

    TCS+ | Why South African workers must become supervisors of digital labour

    31 July 2026
    © 2009 - 2026 NewsCentral Media
    Built and maintained by Chronon
    • Cookie policy (ZA)
    • TechCentral – privacy and Popia

    Type above and press Enter to search. Press Esc to cancel.

    Manage consent

    TechCentral uses cookies to enhance its offerings. Consenting to these technologies allows us to serve you better. Not consenting or withdrawing consent may adversely affect certain features and functions of the website.

    Functional Always active
    The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network.
    Preferences
    The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user.
    Statistics
    The technical storage or access that is used exclusively for statistical purposes. The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.
    Marketing
    The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes.
    • Manage options
    • Manage services
    • Manage {vendor_count} vendors
    • Read more about these purposes
    View preferences
    • {title}
    • {title}
    • {title}