Daily edition October 7, 2026
Daily AI-generated issue
AI engineering, October 7, 2026
Mistral launches Large 4 with top-five cyber scores, OpenAI shares Lean proofs on open math problems, Google ships EmbeddingGemma 2 for on-device multimodal search, TII's 270M OCR beats Opus 5.5
This issue is researched and written by AI models, and every fact is checked against its cited source. No human edits it before it is sent.
- Agents
- APIs
- Audio
- Benchmarks
- Business
- Edition
- Daily
- Published
- October 7, 2026
- Read time
- 3 min
- Stories
- 3
Mistral says Claude Opus 5.5 and GPT-6 Astra score near zero on a real vulnerability-patching test. Not because they fail. Because they refuse. Mistral Large 4 scored 82% on it.
Meanwhile, OpenAI published new results on open problems in mathematics from an internal frontier model, with Lean proofs on GitHub. Google released EmbeddingGemma 2, which maps text, images, audio, video and code into one embedding space for on-device search and RAG.
If you read one thing today: EmbeddingGemma 2. It is Apache 2.0, runs text-only in about 191MB of RAM on a Pixel 11 Pro, and gives you one model for cross-modal retrieval.
Key takeaways
- 01It scores 82% on a test that reproduces and patches a real vulnerability, the highest score reported, plus 93% on Cybench.
- 02On AutomationBench (657 workflows across Gmail, Sheets, Slack and Salesforce) it hits 59.9%, ahead of Kimi K3 and DeepSeek V4 Pro.
- 03For coding it reaches 61.7% on DeepSWE v1.1 and 49.8% on the Coding Agent Index.
- 04It resists 93.3% of prompt-injection attacks on Lakera's B3 benchmark.
- 05API prices are $1.36 per million input tokens and $4.18 per million output tokens.
Top News · 1,668 HN points
Mistral Large 4 is the company's largest and most capable model to date. It is a natively multimodal, open-weight MoE with 1 trillion total and 49 billion active parameters. On the Artificial Analysis Cyber Index, it ranks among the global top five.
Here's what changed:
- It scores 82% on a test that reproduces and patches a real vulnerability, the highest score reported, plus 93% on Cybench.
- On AutomationBench (657 workflows across Gmail, Sheets, Slack and Salesforce) it hits 59.9%, ahead of Kimi K3 and DeepSeek V4 Pro.
- For coding it reaches 61.7% on DeepSWE v1.1 and 49.8% on the Coding Agent Index.
- It resists 93.3% of prompt-injection attacks on Lakera's B3 benchmark.
- API prices are $1.36 per million input tokens and $4.18 per million output tokens.
One catch: this is a preview with the RL run still in progress, and Mistral says weights are due by the end of October 2026.
Try it: the public preview API on Mistral Studio. Also available via Vercel
Sources: mistral.ai · mistral.ai
Top News · 669 HN points
OpenAI is publishing new results on open problems in mathematics. The work comes from an internal frontier model, with no model name, version, or access terms specified.
Here is what makes this credible:
- The proofs come as Lean formalizations readers can inspect.
- OpenAI shares the Lean proof formalizations and research details on GitHub.
- The results come from an internal frontier model applied to open problems.
One catch: OpenAI names no model, version or access terms, so there is nothing to integrate yet.
Why care? Lean formalizations let readers inspect the proofs.
Sources: openai.com
Top News · 275 HN points
EmbeddingGemma 2 maps text, images, audio, video and code into one shared space. It is built for local search, retrieval and RAG.
Here's what you can build with it:
- Its MTEB Code score rises from 68.76 to 78.68 over EmbeddingGemma.
- The full model has 740M parameters, and text-only use needs as little as 270M.
- On a Pixel 11 Pro, quantized RAM is about 191MB for text-only or 567MB multimodal.
- Matryoshka vectors shrink from 768 to 128 dimensions for up to sixfold storage savings.
- An 8K-token context window, four times EmbeddingGemma 1, and Apache 2.0 licensing.
Pair it with Gemma 4 for local RAG, since they share a text tokenizer and audio encoder.
One catch: Google gives no numerical speed results.
Try it: download weights from Hugging Face or Kaggle, deploy with MediaPipe or LiteRT, or run in the browser with transformers.js. It also serves via sentence-transformers, vLLM, llama.cpp and Ollama.
Sources: blog.google · deepmind.google
Also in this issue
Signals
- Google makes Gemini Nano Banana 2.1 generally available with panoramic 8:1 image generation at Flash speed · 23 HN points
- Kandinsky 6.0 Video generates 5-second clips with synchronized 44 kHz audio and lip-sync
- TII's 270M-parameter Falcon OCR Arabic hits 81.9% accuracy, beating Claude Opus 5.5 and GPT Astra
- REA gives coding agents one MCP to reverse engineer native binaries, Electron apps and websites · 279 Stars today
- OpenMontage turns AI coding assistants into video studios with 12 pipelines and 100+ tools · 108 Stars today
- TerrainSR upscales 50x50km heightmaps from 100m to 10m in 0.74s on an RTX 4070