What shipped & dropped across AI labs, today.
Today
Wed · Oct 7 · 1 postYesterday
Tue · Oct 6 · 2 postsEmbeddingGemma 2: The Developer Guide
EmbeddingGemma 2 is a compact, open-source multimodal embedding model that maps text, code, images, video, and audio into a unified 768-dimensional space. Developers can use the sentence-transformers library to selectively load modular modality encoders—ranging from 270M to 740M parameters—to optimize memory usage. Additionally, Matryoshka Representation Learning enables dynamic dimension truncation down to 128d, significantly reducing vector database storage requirements while maintaining high retrieval performance.
Bring multimodal semantic search to the edge with EmbeddingGemma 2
EmbeddingGemma 2 is a new 740M open-weight multimodal model that maps text, images, video, and audio into a unified vector space for privacy-first, on-device retrieval. Developers can easily integrate these capabilities cross-platform using MediaPipe Tasks or optimize fine-grained performance across CPU, GPU, and NPU accelerators with LiteRT. The model enables ultra-low-latency local solutions like search-as-you-type media retrieval, keyframe video moments finding, and zero-shot intent routing.
Monday
Mon · Oct 5 · 1 postFriday
Fri · Oct 2 · 1 postThursday
Thu · Oct 1 · 3 postsEmbeddingGemma 2: an open, lightweight multimodal embedding model
Customize Claude Code with mods
Mods are small TypeScript functions that change how Claude Code works. Rewrite prompts, block risky commands, add custom UI, or replace built-in features. Write one yourself or ask Claude Code to write it, then share it as a plugin.
Barclays scales Claude to upgrade operations and improve client experience
September 30
Wed · Sep 30 · 3 postsClaude for Government is now generally available
Claude Code CLI and Claude for Microsoft 365 also now available in early access.
Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs
To address the quadratic latency bottleneck of self-attention in high-resolution video diffusion models, developers implemented Sparse VideoGen (SVG) to dynamically route attention heads to highly structured spatial or temporal sparse masks. Translating this algorithmic sparsity into physical hardware speedups on TPUs required optimizing the Splash Attention kernel by bypassing empty memory tiles, restricting exact coordinate masking strictly to boundary tiles, and permuting token memory layouts into temporal-major order for contiguous access. By aligning these sparse masks with actual hardware tile execution, the combined optimizations significantly reduced wasted matrix operations and achieved up to a 1.69x end-to-end inference speedup for 1440p video generation.
How Anthropic's sales team rebuilt inbound with Claude Managed Agents
How a Claude-powered buying agent now answers most inbound customers, and how that changed the way our sales team works
September 29
Tue · Sep 29 · 1 postSeptember 28
Mon · Sep 28 · 2 postsTeam Bots: shared AI teammates that learn as they work
Give a Grok Bot the files, apps, and expertise it needs, then share it so your whole team can work from the same context.
Claude Sonnet 5.5 in GitHub Copilot
Claude Sonnet 5.5, Anthropic’s newest Sonnet model, is now generally available in GitHub Copilot. It is designed for well-scoped everyday work like building features and fixing bugs. In our early…
September 25
Fri · Sep 25 · 1 postSeptember 24
Thu · Sep 24 · 1 postSeptember 23
Wed · Sep 23 · 1 postSeptember 22
Tue · Sep 22 · 3 postsOpenAI’s GPT-6 Sol and GPT-6 Luna now available
OpenAI’s GPT-6 family is expanding in GitHub Copilot with two additional models: GPT-6 Sol, and GPT-6 Luna. Joining the previously released GPT-6 Astra, these new options let you select the…
Grok Bot Customer Support
What a task costs on Opus 5.5
The same token price can cost very different amounts per task. Learn what Claude Code tasks cost on Opus 5.5 and which settings change the bill.