TLDR Dev 2026-08-14
Gemini 3.7 Flash ⚡️, understanding is a bottleneck 🧠, how different AI models differ ⚔️
CompactifAI API: get frontier-class AI and cut costs up to 65%. First month free! (Sponsor)
Plug CompactifAI API into your coding agent and cut your bill drastically, without giving up top-tier AI quality. Now with a free first month for new sign-ups.
The best API for output-heavy workloads:
- GLM 5.2 at $3.50 per 1M output tokens: same intelligence tier as Sonnet 5, at 65% lower cost.
- Drop-in for the dev tools you already use like n8n, LiteLLM and Cursor: no migration, no workflow change.
- Try our uncensored versions of GLM 5.1 and Qwen 3.6 27B, ready to answer fully on any topic without giving up accuracy.
- First month free for new sign-ups. No commitment, cancel anytime.
Sign up and start building for free →
Choosing an AI model: one prompt, 11 models, very different results (19 minute read)
A range of AI models, including DeepSeek, Qwen, and Kimi, were tested for generating web pages, showing large variations in performance and credit costs. The best-performing models for a static coffee shop website were DeepSeek V4 Flash, which consumed only 2.4 credits on average, and GPT 5.6 Sol in low effort mode, which averaged 141 credits. In contrast, higher-tier models like Claude Opus 5 showed a lot of credit usage with results that didn't necessarily justify their cost.
How Compaction Works in Pi (6 minute read)
Compaction in coding agents like Pi is a process used to summarize and manage the growing conversation history that can exceed the model's context window during log coding sessions. By creating a smaller representation of prior messages, compaction makes sure that the conversation can continue while preserving relevant information, preventing errors related to context overflow.
Understanding is the new bottleneck (12 minute read)
As AI agents increasingly take on coding tasks, humans still need to understand the code being generated to actively participate in the creative process for collaboration purposes. Effective techniques for improving this understanding include creating structured code explainers, utilizing interactive environments to intuitively grasp systems, and developing shared spaces for team discussions.
Blog about things you don't understand yet (9 minute read)
Blogging serves as a learning tool, where the author consistently gains new insights through the writing process. This approach not only enriches the author's understanding but also offers value to readers, as it forces clear communication and thoughtful engagement with complex topics.
Introducing Gemini 3.7 Flash (11 minute read)
The new Gemini 3.7 Flash model significantly enhances coding and development capabilities by improving accuracy and efficiency in software engineering and web development tasks while offering it at a reduced cost compared to its predecessor. This latest version also incorporates better safeguards, making it more robust for handling complex workflows in knowledge-intensive fields and enhancing the developer experience with improved adaptability and reduced manual oversight.
DeepSeek Harness developer preview: Everything is a plugin (Website)
DeepSeek Harness is now in developer preview, allowing developers to create agent harnesses using a modular plugin system where every capability, such as models, tools, and scheduling, can be swapped or extended. The platform supports multiple runtime modes and ensures complete traceability of agent actions through session logs.
Empty shelves or lost keys? Recall is the bottleneck for parametric factuality (11 minute read)
LLMs such as Gemini 3 and GPT-5 often encounter factual errors not due to a lack of encoded knowledge, but because they struggle with recalling this stored information. By using a knowledge profiling framework, analysis shows that while encoding is nearly saturated in these models, recall is still a hard challenge.
Text AI watermarks will always be trivial to remove (12 minute read)
The European Union AI Act, which takes effect in August, mandates that all AI-generated text be watermarked to make sure it can be identified as artificially generated. Despite proposed solutions like Google's SynthID and potential Unicode tricks from other providers, the methods for watermarking text can still be easily circumvented.
Accelerating GPT-5.6 Sol Ultrafast with OpenAI (5 minute read)
Cerebras and OpenAI have introduced a new service tier, GPT-5.6 Sol Ultrafast, which increases AI processing speeds up to 750 output tokens per second without sacrificing quality. It is powered by Cerebras' Wafer-Scale Engine, which eliminates memory bandwidth bottlenecks by packing 44 GB of SRAM on-chip to keep model weights local. This allows organizations to use AI more effectively in critical scenarios, such as addressing production outages rapidly or responding to cyberattacks.
One Week of Building and Reviewing Code With LLM Agents (34 minute read)
A week-long experiment using LLM agents for backend coding and review revealed insights into productivity, error detection, and the evolving role of human oversight in software development processes.
NP-overrated (3 minute read)
Despite the common belief that NP-hard problems are practically unsolvable, advancements in algorithms and heuristics have demonstrated that efficient solutions can often be found for a majority of cases.
Ordinary Abundance (9 minute read)
Modern conveniences and technologies that once inspired awe have become commonplace in daily life, reflecting humanity's relentless pursuit of progress and the legacy of past innovations.
How (some) Chinese AI Practitioners View Model Distillation (26 minute read)
Chinese AI practitioners are increasingly recognizing the controversial practice of model distillation, which involves using data from superior "teacher" models to improve the capabilities of smaller "student" models, as a technical strategy that is misunderstood and stigmatized.
The most important software engineering news in one daily email
Join 470,000 readers for
one daily email