Data Privacy Wake-Up Call: Lessons From the ZCode Repository Upload Incident

Deep News
Yesterday

Paying for AI coding assistance does not equate to granting permission for the wholesale transfer of an entire code repository. The recent data upload controversy triggered by Z.AI's ZCode feature serves as a critical reminder for every AI user navigating this new landscape. Today's discussion centers on a programmer's code repository, but in a different work context, it could just as easily involve client proposals, interview transcripts, pricing sheets, or confidential business records.

With the rise of office agents and coding agents, where AI evolves from a simple chat window into an assistant capable of reading files and operating computers, the boundaries of data handling cannot afford to be ambiguous, nor can users be left to guess. A clear line must be drawn.

Timeline of the Incident

On September 18th, Z.AI issued an apology concerning the repository upload issue, explaining that the default-enabled Repo Wiki feature could potentially trigger uploads. The company stated the issue had been fixed, that uploaded data was destroyed immediately after page generation, and pledged to open-source the code for third-party review. The original whistleblower, ferstar, subsequently provided an update: a reportedly 313MB commercial project archive had not been successfully uploaded, while a smaller public repository was shown to have been received by the server. Investigations also confirmed that version 3.14.0 had removed the relevant upload mechanism.

In an update dated September 19th, the remediation status was clarified: Z.AI apologized on September 18th at 17:44. Reverse engineering of the latest version 3.14.0 confirmed the physical removal of the repoSnapshot upload pipeline, and the cloud credential endpoint now returns a 404 error. Unresolved issues remain: the official statement claiming "data destroyed immediately after Wiki generation" contradicts the architecture's stated "checkpoint rollback" mechanism. Whether previously uploaded snapshots have been physically purged remains externally unverifiable. To clarify, the 313MB commercial repository snapshot failed to upload 564 times, remaining stuck in a local pending state; OpenWrt traffic logs confirm it never left the local network. A separate 15KB public repository snapshot was indeed received by the server. Z.AI's response focuses on fixing current software behavior, which is a separate matter from accounting for the fate of historical data.

The public's concern revolves around "which of my data was affected." If a vendor only responds with "the feature has been fixed," the two parties are not addressing the same question. Continuous communication and transparency regarding progress are essential for users and the public.

The Core Question: Who Decides for the User?

The fundamental issue that demands scrutiny is: who made the decision on behalf of the user? Cloud-based AI processing typically requires receiving relevant inputs. A user submitting a piece of code for error analysis has a relatively clear expectation of that data transfer. However, reading the content required for the current task versus packaging up the entire repository and its history involves different scopes of data and should warrant clear disclosure and consent.

Opening a project should not imply consent to upload all its materials. Nor should the potential for an improved user experience justify a default assumption that users are willing to surrender more data. Product design cannot mistake "convenient implementation" for "implied authorization."

It is particularly alarming that users often cannot understand what a privacy toggle actually controls. Does disabling "experience optimization" mean refusing training, refusing telemetry, or refusing file transfers? If these behaviors are governed by separate mechanisms, the interface must explain them clearly. Requiring ordinary users to dissect a client to discover where files are sent is, in itself, a failure of product transparency. Four distinct questions cannot be conflated into a single answer.

Data Handling: Separate the Stages

Upload, retention, training, sharing, and sale are distinct stages. Discovering an upload does not directly infer training or sale; a vendor's promise not to train does not automatically assure users there has been no upload or retention. The current Chinese privacy policy for ZCode states that the optimization program is off by default and that data will not be used for product and model training or optimization before users actively opt in. This commitment should be accurately communicated, and its execution needs to be verified against actual behavior.

Similarly, claims of "encryption" address how data is protected but cannot replace user consent. Users need to know who holds decryption capabilities and who has access permissions. "Deletion after use" requires clarification on which systems deletion occurs, whether backups, caches, and derivative content are involved. A vague security statement cannot answer these specific questions.

Don't Deflect with Distractions

The motives of Taiyuan Chengming Technology Co., Ltd., which sent a formal inquiry to Z.AI, have become a topic of speculation. Public reports relay the company's demands for accountability, but currently there is insufficient evidence to prove it deliberately staged the situation or falsified evidence. Claims regarding which data was affected and the extent of the losses still require evidentiary support and cannot be fully accepted based solely on a letter. Even if someone intentionally tests software, as long as there is no tampering with the client, falsifying logs, or explicit authorization of the disputed upload, the testing motive cannot absolve the product's default behavior. Whether genuine losses exist can continue to be verified; whether default actions exceeded user expectations should also be independently reviewed. These two matters do not have to be bundled together.

Restoring Trust Requires Verifiable Results

To rebuild trust, vendors should, at minimum, disclose the affected versions and time ranges, allow users to query which data was uploaded, explain the recipients, purposes, retention periods, and the outcome of historical data processing, and ensure reviews cover the controversial older versions and server-side processes. Z.AI's pledge to open-source its code is welcome, but releasing only the patched client cannot independently prove that cloud historical data has been deleted.

Products should also offer genuinely understandable and actionable choices: optional full-repository uploads should be off by default, data scope should be explained before transfer, options to exclude sensitive directories should be provided, and users who decline additional data usage should still be able to access basic services. Security should not only be written into lengthy agreements but reflected in every default option.

For enterprises procuring AI services, evaluation checklists should include additional criteria. Beyond generation quality, speed, and price, they must ask about data flow, deletion mechanisms, access permissions, and anomaly notification procedures. A tool should not be allowed to operate outside internal data management rules simply because it started as a personal trial. Clear rules should be established in advance regarding whether a company permits the upload of certain types of materials.

Users, too, need to adjust their habits: handle sensitive tasks in isolated environments, grant access only to necessary directories, anonymize client data, rotate active keys that may be exposed, and preserve versions, timestamps, and logs when anomalies are detected, rather than deleting evidence before a full investigation. These measures do not absolve vendors of responsibility, but they can reduce passive reliance on a mere promise.

As AI becomes more capable, authorization must become more specific. We all want cheaper, better tools, but we also need to know where our data is at all times. Verification conducted as of September 20, 2026. This article is based on public disclosures, policies, and reports; it did not independently run the affected client for reproduction, nor does it treat all claims from the inquiring company as established facts.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10