Switch language한국어
Back to the list

Prompt injection disclosures: 4 labs compared

TL;DR AI

Key summary

2 min read
  1. A comparison of Anthropic, OpenAI, Google, and Meta shows there is still no common standard for reporting prompt-injection risk.

  2. Anthropic released the most detailed system card, breaking out four attack surfaces and reporting a 31.5% browser attack success rate before safeguards.

  3. OpenAI gave a single connectors robustness score, Google folded the issue into a broader safety framework, and Meta did not publish a closed-model card.

  4. Security experts say the inconsistent disclosures make it difficult for buyers and security teams to compare models or set a shared benchmark.

Read the original