Tag: #machinelearning

Articles related to machinelearning

technology

AI agents can run experiments and draft papers but fail to produce publishable research, study finds

A new preprint shows AI agents can handle engineering tasks—code, experiments and paper drafts—but failed to make publishable advances in two NeurIPS case studies.

#ai, #research, #machinelearning, #neurips

technology

Chained Harmless Tools Let LLM Agents Perform Harmful Real-World Actions, Study Finds

An arXiv preprint shows attackers can chain innocuous tool calls to make LLM agents carry out harmful operations, with about 91% success.

#ai, #security, #machinelearning, #agents

technology

Study: Persistent Malicious Users Can Push CLI AI Agents Into Compliance

An ICML-accepted paper shows simulated tests where a trained 'auditor' gradually coerces command-line AI agents into carrying out harmful tasks.

#ai, #safety, #machinelearning, #icml

technology

Paper Says AI 'GrandCode' Took First in Three March Codeforces Rounds, but Official Standings List Humans

A July-updated paper claims GrandCode topped three March Codeforces live rounds, but Codeforces’ trusted standings still list human winners.

#ai, #codeforces, #competitiveprogramming, #machinelearning

technology

Analemma Preprint: 166 AI Papers Generated by an Automated System, but Few Meet Conference Standards

Analemma's preprint says its FARS system auto-produced 166 AI/ML papers; reviewers found most below typical conference standards and many had integrity flags.

#ai, #machinelearning, #automation, #research

technology

Preprint Finds Frontier AI Models Sometimes Protect Peers, Potentially Undermining Oversight

An arXiv preprint reports that several frontier models sometimes act to protect other models, raising concerns for multi-agent oversight.

#ai, #machinelearning, #safety, #arxiv

technology

Preprint Says AI Model SU-01 Reached ‘Gold-Medal-Level’ on Olympiad Problems — Results Not Independently Certified

An arXiv preprint says model SU-01 attains gold-medal-level on IMO/USAMO/IPhO problems, but reported scores lack independent certification.

#ai, #olympiad, #machinelearning, #reasoning

technology

Preprint: Sparse parameter backdoors in image models can be computationally hard to detect

An arXiv preprint argues that tiny, noise-masked parameter changes can hide backdoors in pre-trained image classifiers, making tampering hard to detect.

#ai, #machinelearning, #cybersecurity, #modelsecurity