Execution Is Cheap Now — Evals Are the Job
I routed my whole workday through Claude — email, Slack, research, tickets. The throughput gain was the boring part. The interesting part is what's left for me to do.
17 posts
I routed my whole workday through Claude — email, Slack, research, tickets. The throughput gain was the boring part. The interesting part is what's left for me to do.
A casual LINE debate over whether restaurants need stablecoin payments turns into a bigger question: is this an era where B2C startups get flattened by giants?
A conversation with a Silicon Valley product manager — the traffic-light framework for PM decisions, whether to validate a side project with interviews or a POC, the real inside story on DeepMind, Anthropic, OpenAI and Meta, and how hard it is to find a mentor in Silicon Valley.
Notes from a walk with Cursor's Ian Huang in San Francisco: coding agents, product validation and startups, then personal brand, education, philosophy, and how to build the life you actually want.
Hugging Face solves a different problem than GitHub. How ML engineers use both platforms, and why models and datasets live on Hugging Face.
See how Claude Code, LangGraph, and n8n split GTM engineering work — operator harness, reasoning engine, and automation glue — with real workflow examples.
Hugging Face started as a chatbot for teenagers. Here's the week it pivoted to open-source AI infrastructure — and what that pivot teaches builders.
Clay's GTM engineering stack shows the difference between production agent frameworks like LangGraph and operator harnesses like Claude Code or Codex.
GTM engineering turned go-to-market from a series of manual handoffs into a system you build. A look at what the role actually does, how it differs from RevOps, and why ~100 job listings for it now open every month.
Paying for Linear or Jira might be one of the worst investments a company makes. A first-principles look at why PM tools exist, and what replaces them once AI can render the interface.
DESIGN.md is Google's open file format for handing your entire design system to an AI coding agent — like a CLAUDE.md, but for design. Here's what's in it, what Atlassian found when they stress-tested it, and why it's a near-perfect fit for a small site like this one.
Ramp built Inspect, Stripe built Minions, Block built goose. Cursor and Claude Code exist and are excellent — so why are these companies rolling their own? The answer is context, integration, and ownership, and it says something about where the moat in AI engineering actually sits.
Notes from an interview with Perplexity's Dmitry Shevelenko — why every knowledge worker is becoming an executive, what Perplexity learned watching CEOs out-adopt their own employees, and why agentic checkout is overhyped.
Notes from a coffee chat with David Zeng, co-founder of Beacons AI — on cancelling AI subscriptions, why the boring jobs are the ones worth paying for, and what 'Service as a Software' really means.
A plain-English guide to the three features that decide what an AI agent can do — MCP (access), Skills (a toolbox it chooses from), and Workflows (a strict assembly line) — with the analogy that finally made them click for me.
Notes from the Crossroads podcast on Forward-Deployed Engineers (FDEs), why most enterprise AI projects fail, and the shift from selling slide decks to shipping working agents.
Anthropic says 80% of its own code is now written by Claude. I read their recursive self-improvement essay, checked the numbers, and separated what they actually claimed from the doom-spin around it.