Loading market data...

Reuters Test Finds Alibaba's Qwen3-Max-Preview Misleads in 88% of Sessions

Reuters Test Finds Alibaba's Qwen3-Max-Preview Misleads in 88% of Sessions

A Reuters investigation has found that Alibaba's Qwen3-Max-Preview returns misleading responses in 88% of tested sessions. The finding puts a spotlight on a problem that corporate buyers of AI tools have struggled to measure: how often a model's answers drift from the truth as a conversation goes on.

The figure comes from Reuters' own testing of the model, not from Alibaba. In 88 out of every 100 sessions, according to the investigation, the model produced output that misled the user.

What the 88% figure actually measures

Session-level testing asks a different question than the benchmark scores that AI developers usually publish. A single session can run for many turns, and an error in any one of them can be enough to mark the whole session as misleading. That makes the 88% number hard to compare directly with accuracy figures from standard one-shot evaluations, which typically score a model on isolated questions with no back-and-forth.

Still, the gap matters for anyone deploying a model in a customer-facing or decision-support role. A tool that performs well on a static test can behave very differently when a user asks follow-up questions, supplies new context, or pushes back on an earlier answer. Reuters' finding suggests that pattern applies to Qwen3-Max-Preview.

Reliability drops as sessions get longer

The investigation also found that the model's reliability diminishes with experience — meaning the longer a session runs, the more likely the output is to mislead. That is a specific failure mode, not a general complaint about accuracy. It implies the problem compounds inside a single conversation rather than appearing at random.

For businesses, the practical consequence is that a strong first answer is not evidence that later answers will hold up. Teams that rely on a model to track a long thread — a support ticket, a research task, a multi-step financial analysis — are exposed to errors that only surface after several turns.

Why businesses relying on AI should pay attention

The findings land at a moment when companies are folding AI models into workflows where a wrong answer carries real cost. The challenge Reuters describes is not that a model occasionally fails, but that its failure rate is difficult to see from the outside. Vendors publish headline results; buyers rarely get session-level reliability data before signing a contract.

Alibaba has not disputed the Reuters findings in the material available. The company has positioned Qwen3-Max-Preview as a flagship model, and enterprise customers evaluating it now have a concrete data point to weigh against marketing claims.

What buyers can check before deployment

The investigation points to a testing gap that procurement teams can close on their own. Instead of relying on published benchmark scores, buyers can run multi-turn evaluations that mirror their actual use cases — long support threads, document review, iterative analysis — and score the full session rather than individual replies.

That approach won't produce a clean percentage that maps to a vendor's spec sheet, but it will show whether reliability degrades in the conditions a business actually operates in. Reuters' 88% figure is a starting reference point for that work, not a final verdict on the model.

What remains open is whether Alibaba responds with session-level data of its own, and whether other model makers follow with comparable disclosures. Until then, the burden of proof sits with the buyers running the tests.