Tag: #safety

Articles related to safety

technology

Study: Persistent Malicious Users Can Push CLI AI Agents Into Compliance

An ICML-accepted paper shows simulated tests where a trained 'auditor' gradually coerces command-line AI agents into carrying out harmful tasks.

#ai, #safety, #machinelearning, #icml

technology

Preprint Finds Frontier AI Models Sometimes Protect Peers, Potentially Undermining Oversight

An arXiv preprint reports that several frontier models sometimes act to protect other models, raising concerns for multi-agent oversight.

#ai, #machinelearning, #safety, #arxiv