ARCHIVE
아카이브
기존에 공개한 글을 원래 주소 그대로 보관합니다.
과거의 글에는 작성 당시의 기술과 관점이 담겨 있습니다.
356개 글 · 7 / 12 페이지
· EN
Context Engineering: The Core Skill Behind Production AI Agents
Why context engineering has become the defining skill for production AI agents in 2026 — 4 critical failure patterns and 5 core techniques, from an Engineering Manager perspective.
· EN
Karpathy's autoresearch: 100 Autonomous ML Experiments Overnight
Andrej Karpathy's autoresearch is a 630-line open-source tool that lets AI agents autonomously iterate ML experiments overnight. We analyze R&D team adoption strategies from an EM perspective.
· EN
LLMs Unmasking Anonymity — Reality of Large-Scale Identity Tracking
Analyzing large-scale online deanonymization research using LLMs and presenting organizational security defense strategies for engineering leaders.
· EN
AI Reliability Engineer and the Centaur Pod Model in 2026
Junior roles are evolving into AI Reliability Engineers. Centaur Pod team structures, Code Audit hiring, Defect Capture Rate — the AI-native team design brief for Engineering Managers.
· EN
Claude Found 22 CVEs in Firefox — AI Security Audits Arrive
Anthropic Claude Opus 4.6 discovered 22 CVEs in Firefox in just two weeks. We break down how AI-driven security audits work and what engineering leaders should do next.
· EN
Agent Scaling Science — Google Debunks 'More Agents = Better'
Google Research's 180-configuration experiment exposes the multi-agent paradox: 39–70% degradation on sequential tasks, 17.2× error amplification, and what it means for your architecture.
· EN
RoguePilot — Copilot Prompt Injection and AI Tool Security
Analysis of the RoguePilot vulnerability found in GitHub Codespaces, passive prompt injection risks in AI coding tools, and security guidelines for engineering teams.
· EN
A2A + MCP Hybrid Architecture: 2026 Multi-Agent Production Strategy
Google A2A and Anthropic MCP are complementary, not competing. An EM/CTO view of the two protocols' roles and strategies for running multi-agent systems safely in production.
· EN
Cursor Agent Trace — An Open Standard for Tracking AI-Generated Code
Analyze Cursor Agent Trace 0.1.0 specification and discover why AI code attribution tracking is critical for engineering leaders and CTOs beyond git blame.
· EN
Cut Agent Fleet Costs by 90% with Heterogeneous LLM Architecture
The Plan-Execute pattern: large models plan, small models execute. A practical guide for EMs and CTOs on heterogeneous LLM architecture strategies to dramatically reduce agent fleet costs without sacrificing quality.
· EN
Tool-R0: Self-Play RL for Tool-Using Agents, Zero Data
The arXiv paper Tool-R0 achieves 92.5% improvement in LLM tool-calling via Self-Play RL alone, with no training data. We analyze its Generator-Solver co-evolution and practical implications.
· EN
Bayesian Teaching: How LLMs Learn Probabilistic Reasoning
Google's Bayesian Teaching research, published in Nature Communications, introduces a training methodology that enables LLMs to probabilistically update their beliefs when receiving new information.
· EN
Deloitte 2026 Agentic AI: Why 89% Fail and the EM Framework
Only 11% of enterprises run Agentic AI in production. The barrier isn't technology—it's operational model. Here's the Delegate-Review-Own framework for EM/VPoE.
· EN
MCP Security Crisis — 30 CVEs in 60 Days
The MCP attack surface is expanding fast. Here's an analysis of 30 CVEs, a three-layer attack model, and an enterprise security hardening checklist for engineering leaders.
· EN
ADL: The OpenAPI of AI Agent Governance
ADL declaratively defines AI agent roles, permissions, and allowed tools—the OpenAPI of agent governance. Covers the core spec structure and practical EM/CTO governance adoption strategies.
· EN
Cognitive Debt: The Liability AI Teams Quietly Accumulate
Anthropic's 2026 Agentic Coding Trends Report heralds a productivity revolution, while parallel research warns of Cognitive Debt: as AI writes more code, teams quietly lose shared understanding.
· EN
How to Build an Elite AI Engineering Organization in 2026
Analysis of the elite AI engineering culture that topped Hacker News. Understanding the 5.7x gap between $3.48M vs $610K revenue per employee, and the Taste × Discipline × Leverage formula every EM should practice
· EN
Olmo Hybrid — 2x Data Efficiency with Transformer + RNN
AI2's Olmo Hybrid pairs Transformer and DeltaNet in a 3:1 mix for 49% token savings at equal accuracy. Architecture deep-dive and LLM engineering implications.
· EN
AI's Convenience Loop: How Coding Tools Reshape Language Choices
AI coding tools create a convenience loop reshaping language popularity. Why TypeScript surged 66% and Python hit #1, with EM/CTO tech stack decision framework.
· EN
Meta Llama 4 Deep Dive — Maverick & Scout for Enterprise
A technical and strategic analysis of Meta Llama 4 Maverick (400B MoE) and Scout (10M context): architecture, benchmarks, cost structure, and what it means for open-source AI strategy.
· EN
NIST AI Agent Security Standards
Understanding NIST AI Agent Standards Initiative and an actionable security checklist for Engineering Managers to strengthen AI agent security within their teams.
· EN
Optimizing AI Agent Workflows with Meta-Tools: AWO Framework Guide
Analyze the Agent Workflow Optimization (AWO) framework from arXiv research. Compile repetitive tool call patterns into meta-tools to reduce LLM calls by 12% and improve success rates by 4%.
· EN
Claude Cowork Enterprise Launch: From Dev Tool to Platform
Analysis of Anthropic Claude Cowork enterprise features. Plugin Marketplace, MCP connectors, Excel and PowerPoint integration — a CTO strategy guide for org-wide AI adoption.
· EN
Deep-Thinking Ratio: Cut LLM Inference Costs by 50%
Google & UVA research overturns the "longer = better" assumption for LLM reasoning. The Deep-Thinking Ratio (DTR) can cut inference costs in half while improving accuracy.
· EN
MCP Joins the Linux Foundation
Anthropic donated MCP to the Linux Foundation, with OpenAI, Google, and Microsoft on board. With 76% of companies exploring adoption, here is a practical strategy guide for EMs and VPoEs.
· EN
MIT EnCompass: +40% AI Agent Accuracy via Search
Discover how MIT CSAIL's EnCompass framework applies search strategies to AI agent execution paths, dramatically improving reliability and accuracy in production.
· EN
Jira AI Agents and MCP — What Engineering Managers Need to Know
Atlassian has officially launched AI agents in Jira and adopted MCP platform-wide. Here's what engineering managers need to prepare for organizational change.
· EN
LLM Coding Harness Optimization
LLM coding harness (edit formats, tool interfaces) beats model swaps with 5-14% gains. Grok Code Fast: 6.7%→68.3%. Harness engineering guide for EMs and CTOs.
· EN
AI Model Distillation Attacks — IP Protection for CTOs
Analyzing Anthropic's detection of large-scale AI model distillation attacks and presenting practical strategies for enterprises to protect intellectual property when using AI APIs.
· EN
Anthropic vs Pentagon — CTO Vendor Strategy for the AI Governance Era
Analyzing Anthropic's refusal of Pentagon military AI demands and providing practical guidance for CTOs/VPoEs on establishing AI vendor dependency risk and governance strategies.