Anthropic Silently A/B Tests Claude Code, Degrading User Workflow & Transparency | Evidence Inside

Do Not A/B Test My Workflow

The promise of artificial intelligence lies in its ability to augment human capabilities, to streamline complex tasks, and to provide tools that empower professionals. But what happens when the very tools we rely on begin to change, subtly and without our knowledge, impacting our productivity and eroding trust? A growing number of users of Anthropic’s Claude Code are voicing concerns that the company is conducting opaque A/B tests that are actively degrading their experience, turning a professional tool into an unpredictable experiment. This raises critical questions about transparency, user agency, and the responsible deployment of AI technologies.

Claude Code, a powerful AI assistant geared towards software development and complex problem-solving, is a subscription-based service. Users, like many professionals, expect a degree of stability and predictability from the tools they pay for. The recent discovery of hidden A/B tests, controlling core functionality without user consent, has sparked outrage and a demand for greater control and visibility. The core issue isn’t simply that changes are happening, but *how* they are happening – silently, without notification, and seemingly at the expense of user experience. This situation highlights a fundamental tension between a company’s desire to optimize its product and a user’s right to a stable and predictable workflow.

The concerns center around a feature called “plan mode,” used for outlining and structuring complex tasks. Users have reported unexpected changes in how plans are generated, ranging from limitations on length to the removal of crucial contextual information. These alterations aren’t announced; they simply *happen*, leaving users scrambling to understand why their established workflows are suddenly broken. The discovery that these changes are controlled by A/B tests, managed through a system called GrowthBook, and assigned to users without their knowledge, has fueled accusations of a lack of transparency reminiscent of practices seen at companies like Meta, where silent experimentation on users has been a point of contention.

Uncovering the ‘Tengu’ Tests

The issue came to light when a user, going by the handle “backnotprop” on platforms like GitHub and X (formerly Twitter), began investigating unexplained regressions in Claude Code’s plan mode. Through a detailed analysis of the application’s binary code, they uncovered the existence of an A/B test named “tengu_pewter_ledger.” This test controls how plan mode writes its final plan, with four distinct variants: “null,” “trim,” “cut,” and “cap.” Each variant progressively restricts the output, limiting context, explanation, and even the overall length of the generated plan. The findings were documented in a public gist, providing a technical breakdown of the A/B testing framework.

A visual representation of the A/B test variant groups found within the Claude Code binary, as documented by backnotprop.

According to the analysis, the most restrictive variant, “cap,” imposes significant limitations. It caps plans at 40 lines, prohibits contextual background information, and forbids prose paragraphs, instructing the model to “delete prose, not file paths” when exceeding the line limit. Users assigned to this variant reported a drastically different experience, receiving terse bullet-point lists instead of the detailed, conversational plans they were accustomed to. There was no opportunity for back-and-forth dialogue, no ability to steer the AI’s reasoning – simply a pre-determined output.

The cap variant instructions
The specific instructions for the “cap” variant, demonstrating the restrictive parameters imposed on plan generation.

The user who discovered the tests reported receiving the “cap” variant without any prior notification or opt-in. Telemetry data collected at the end of the plan execution logs the assigned variant, revealing that Anthropic is tracking user interactions and gathering data on the performance of each variant. The purpose of this data collection remains unclear, but it underscores the fact that paying customers are, in effect, participating in an experiment without their explicit consent.

Example plan output under the cap variant
An example of a plan generated under the “cap” variant, showcasing the terse bullet-point format and lack of contextual information. Rendered with Plannotator.

The Broader Implications for AI Tooling

This situation with Claude Code raises broader concerns about the ethics of A/B testing in professional AI tools. While A/B testing is a common practice in software development, its application to tools used for critical work requires careful consideration. Unlike social media platforms where subtle changes in algorithms might affect engagement, alterations to AI-powered professional tools can directly impact a user’s ability to perform their job effectively. The lack of transparency and control in this case is particularly troubling, as it undermines the trust that users place in these tools.

Anthropic, founded by former OpenAI researchers, has positioned itself as a leader in AI safety and responsible development. The company’s stated mission is to build AI systems that are beneficial to humanity. However, this incident casts a shadow on that commitment, suggesting a disconnect between the company’s stated values and its operational practices. The concerns echo criticisms leveled against companies like Meta, which have been accused of prioritizing engagement metrics over user well-being through similar silent experimentation tactics.

The recent release of Claude Code Skills 2.0, which includes features like automatic evaluations, benchmarks, and A/B testing, demonstrates Anthropic’s commitment to improving its AI tools. According to Pasquale Pillitteri, this update aims to address challenges like skill obsolescence and unreliable evaluation methods. However, the value of these improvements is diminished if users are unaware of the underlying changes and lack the ability to control their experience. The framework categorizes skills into “Capability Uplift Skills” and “Workflow/Preference Skills,” but the lack of transparency regarding A/B testing applies to both.

The incident has sparked a debate within the AI community about the necessitate for greater user agency and control over AI tools. Many argue that users should have the option to opt-out of A/B tests, to receive clear notifications about changes, and to understand the rationale behind those changes. The ability to configure AI tools to meet specific needs and preferences is also seen as crucial, allowing users to tailor the technology to their individual workflows.

What’s Next?

As of March 14, 2026, Anthropic has not issued a public statement addressing the concerns raised by users regarding the silent A/B tests in Claude Code. However, the issue has gained traction on platforms like GitHub and X, prompting a wider discussion about transparency and user control in AI tooling. Users have filed issues on the Claude Code GitHub repository, requesting greater visibility into the A/B testing process and the ability to opt-out of experiments. One such issue details the concerns and requests for configuration options.

The company is expected to address these concerns in a future update, potentially offering users more control over their experience and providing greater transparency into the A/B testing process. The outcome of this situation will likely set a precedent for how other AI companies approach A/B testing and user agency in the development of professional tools. It remains to be seen whether Anthropic will prioritize optimization at the expense of user trust, or whether it will embrace a more transparent and user-centric approach to AI development.

The debate surrounding Claude Code serves as a crucial reminder that the development of AI is not merely a technical challenge, but also an ethical one. As AI tools become increasingly integrated into our professional lives, it is essential that we prioritize transparency, user agency, and responsible deployment to ensure that these technologies truly empower us, rather than control us.

Do you have experience with unexpected changes in AI tools? Share your thoughts and experiences in the comments below. And please share this article with your network to support raise awareness about the importance of transparency and user control in the age of AI.

Leave a Comment