Training an LLM on CPU in 1.2 Hours
How FlashLM v3 trained a 13.6M-parameter LLM on CPU alone in 1.2 hours with MatMul-Free ternary-weight architecture, and its implications for edge AI.
Tags
2 posts
How FlashLM v3 trained a 13.6M-parameter LLM on CPU alone in 1.2 hours with MatMul-Free ternary-weight architecture, and its implications for edge AI.
Hidden prompt injection telling AI reviewers what to write was found in ICML submission PDFs. We analyze the attack and the risks of AI-dependent peer review.