ARCHIVE
아카이브
기존에 공개한 글을 원래 주소 그대로 보관합니다.
과거의 글에는 작성 당시의 기술과 관점이 담겨 있습니다.
356개 글 · 8 / 12 페이지
· EN
GitHub Agentic Workflows—AI Agents Join CI/CD
Analyzing GitHub Agentic Workflows technical preview. Define automation in Markdown, and AI agents perform issue triage, code reviews, and test generation in Continuous AI paradigm.
· EN
MIT TLT: Doubling Reasoning LLM Training Speed
MIT researchers introduced TLT, accelerating reasoning LLM RL training by 70-210% through adaptive drafters and speculative decoding. Reduces training costs without additional hardware.
· EN
Claude Code Remote Control Complete Guide
A complete guide to setting up and using Claude Code Remote Control. Learn how to monitor and control desktop tasks from your phone with practical workflow examples.
· EN
Switching OpenClaw to OpenAI Codex
OpenClaw migration guide: switch from Claude/Gemini OAuth to OpenAI Codex in 15 minutes. Covers backup, model config, per-agent settings, provider layer strategy, and cost comparison.
· EN
Don't Trust the Salt — Multilingual LLM Guardrail Gaps
An analysis of how LLM guardrails fail in multilingual environments. We examine the structural issues causing safety verification failures in non-English languages and practical countermeasures.
· EN
GGML/llama.cpp Joins Hugging Face
The ggml.ai team joins Hugging Face to secure the long-term sustainability of llama.cpp. We analyze the structural changes and technical implications for the local AI inference ecosystem.
· EN
ASIC Inference Chip Runs Llama 3.1 8B at 16,000 tok/s
Startup Taalas achieves 16,000 tok/s on Llama 3.1 8B using custom ASIC chips without GPUs. We analyze the shift away from GPU dependency and the inference cost revolution.
· EN
Consistency Diffusion Language Models
Together AI introduces CDLM, boosting diffusion language model inference up to 14x faster while maintaining quality. Block-wise parallel generation with KV caching is the key breakthrough.
· EN
Gemini 3.1 Pro Release — Performance Analysis and Claude Comparison
Google releases Gemini 3.1 Pro with 77.1% on ARC-AGI-2, doubling reasoning performance. We analyze benchmarks, compare with Claude, and explore multimodal evolution.
· EN
IQ*_K/IQ*_KS Quantization Merged into llama.cpp
IQ-series quantization methods developed in ik_llama.cpp are being merged into llama.cpp mainline. Learn about IQ2_K through IQ4_KS precision improvements and local LLM inference optimization.
· EN
Qwen3 Coder Next llama.cpp Graph Optimization
ggerganov restructures the llama.cpp compute graph to achieve up to 38% inference speedup for the Qwen3 Coder Next 80B model. Detailed benchmark analysis and technical breakdown.
· EN
DDR5 RDIMM vs RTX 3090 — The Cost-per-GB Tipping Point for Local LLMs
DDR5 RDIMM pricing has dropped below RTX 3090 VRAM per GB, marking a turning point in local LLM hardware decisions. We analyze CPU vs GPU inference cost structures.
· EN
Devstral Small 2 24B & Qwen3 Coder 30B
Mistral Devstral Small 2 24B and Qwen3 Coder 30B arrive simultaneously. A comparative analysis of small coding models that run on Raspberry Pi and the future of local AI coding.
· EN
Kitten TTS V0.8: Sub-25MB SOTA Speech for Edge Devices
A deep dive into Kitten TTS V0.8 — a 14M parameter, sub-25MB text-to-speech model matching cloud TTS quality. Analysis of edge deployment potential and the local voice AI trend.
· EN
$30 Radio + Local AI = Internet-Free Smart Home
Analyzing a real-world project that achieves voice control and smart home automation without internet using just a Mac mini and a $30 LoRa radio. A deep dive into local AI × IoT implementation and costs.
· EN
BarraCUDA: Open-Source Compiler Running CUDA on AMD GPUs
BarraCUDA compiles CUDA to AMD GPU binary — no LLVM or HIP. 15,000 lines of C99 cover shared memory, atomics, warp intrinsics. A direct challenge to NVIDIA GPU vendor lock-in.
· EN
Claude Sonnet 4.6: Anthropic's Mid-Tier Model Strategy
Claude Sonnet 4.6 analysis: model versioning strategy, performance vs Opus and Haiku, API cost changes, and upgrade guidance for developers building on Claude.
· EN
DeepSeek V4 Imminent — China's Next-Gen AI Race Speeds Up
As DeepSeek V4 approaches, Qwen3.5 and GLM-5 keep pace. Reasoning gains over R1, benchmark comparisons, and how open LLMs nearing GPT-4 reshape the global AI landscape.
· EN
KaniTTS2 — Open 400M TTS with Voice Cloning on 3GB VRAM
KaniTTS2 is a 400M-parameter open-source TTS model that runs voice cloning on just 3GB VRAM with full pretraining code released.
· EN
Training an LLM on CPU in 1.2 Hours
How FlashLM v3 trained a 13.6M-parameter LLM on CPU alone in 1.2 hours with MatMul-Free ternary-weight architecture, and its implications for edge AI.
· EN
Does AGENTS.md Actually Work? The First Empirical Study
The first empirical study evaluating AGENTS.md effectiveness has been published. We analyze its impact on coding agent success rates and inference costs.
· EN
AI Self-Generated Skills Are Useless
SkillsBench proves AI agents cannot author useful skills for themselves. Across 7,308 trajectories, self-generated skills showed zero benefit while human-curated skills improved performance by 16.2pp.
· EN
FunctionGemma 270M — 90-97% Tool Calling in a Tiny Model
Analysis of how fine-tuning FunctionGemma 270M improved multi-turn tool calling accuracy from 10-39% to 90-97%, matching a 120B teacher model. More evidence that scaling isn't everything.
· EN
4 of OpenRouter's Top 5 Models Are Open Source
Four of the top five most-used models on OpenRouter are open source (Qwen3-Coder, DeepSeek R2, MiniMax M2.5, etc.).
· EN
Qwen 3.5 Goes Bankrupt on Vending-Bench 2
Qwen 3.5, a top performer on standard benchmarks, goes bankrupt on Vending-Bench 2's vending machine simulation. Exploring the blind spots of benchmark-driven AI evaluation.
· EN
Claude Code with Local Models Triggers Full Prompt Reprocessing
Analyzing the full prompt reprocessing issue when running Claude Code with local LLMs. Learn about KV cache invalidation mechanics and developer tool design lessons.
· EN
Heretic 1.2: 70% VRAM Reduction via Quantization and MPOA Explained
Heretic 1.2 is here with 4-bit quantization cutting VRAM usage by up to 70% and MPOA delivering higher-quality abliteration. A deep dive into the latest cost-saving techniques for local LLM operations.
· EN
Karpathy: AI Training Costs Drop 40% Per Year
Karpathy's analysis reveals AI model training costs fall 40% annually. We examine the structural factors — hardware evolution, algorithm efficiency, and data pipeline optimization — and their industry impact.
· EN
How to Run Qwen3-Coder-Next 80B on 8GB VRAM
Analyzing quantization and lazy loading techniques to run an 80B parameter coding AI model on consumer 8GB VRAM GPUs. Exploring the practicality and limitations of local LLM coding.
· EN
Building SQLite with an AI Swarm
Six AI agents (Claude, Codex, Gemini) built a 19,000-line Rust SQLite clone in parallel. Analyzing the real costs of multi-agent coordination and task division.