Anthropic hiện đang nhúng một mẫu thống kê ẩn—SynthID-Text—vào mọi phản hồi của Claude. Công ty cho biết động thái này nhằm đáp ứng Bộ Quy tắc Minh bạch của Đạo luật AI EU và cung cấp cho các nhà phát triển một cách để chứng minh một đoạn văn bản đến từ mô hình AI. Dấu chìm này vẫn tồn tại sau các chỉnh sửa thông thường mà không làm thay đổi lời văn đối với người đọc.

Tại sao dấu chìm lại được thêm vào

Các cơ quan quản lý châu Âu đã gây áp lực lên các nhà cung cấp mô hình ngôn ngữ lớn (LLM) để làm cho nội dung do AI tạo ra có thể phân biệt được với văn bản do con người viết. Bộ Quy tắc Minh bạch của Đạo luật AI EU, sắp được thực

Critics argue that a hidden signature, even if invisible, could be repurposed for tracking or attribution beyond regulatory compliance. They also fear false positives: a detection API that mislabels human-written text could undermine trust in platforms that rely on the signal. Anthropic acknowledges these risks and says the detection algorithm will be open-source, allowing independent audits and calibration.

Another practical worry is the computational overhead of generating the watermark. Anthropic reports that the additional processing adds less than a fraction of a percent to inference time, a claim that third-party benchmarks will soon test.

What to watch next

  • API rollout – Anthropic plans to release a public endpoint that extracts the hidden bit pattern from any Claude output. Developers will be able to integrate the check into content-moderation pipelines.
  • Standardization efforts – Regulatory bodies and industry groups are drafting a common format for watermark metadata, which could make cross-model detection easier.
  • Empirical studies – Independent researchers are beginning to probe the resilience of SynthID-Text against adversarial rewriting tools. Their findings will shape how much confidence platforms place in the signal.

Bottom line

Anthropic’s adoption of SynthID-Text gives the AI community a concrete tool to meet upcoming European transparency rules while keeping Claude’s output quality intact. The watermark survives everyday proofreading but falls away when text is fully regenerated, and it lives in the non-code parts of programming output where it does not interfere with functionality. Whether the approach becomes the industry norm hinges on how well the detection API performs in real-world settings and whether concerns about privacy and false alarms can be addressed.