stanford-crfm-website
-
1
Advancing Customizable Benchmarking in HELM via Unitxt Integration
-
2
HELM Safety: Towards Standardized Safety Evaluations of Language Models
-
3
General-Purpose AI Needs Coordinated Flaw Reporting
-
4
HELM Capabilities: Evaluating LMs Capability by Capability
-
5
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
-
6
Surprisingly Fast AI-Generated Kernels We Didn’t Mean to Publish (Yet)
-
7
Reliable and Efficient Amortized Model-Based Evaluation
-
8
HELM Long Context
-
9
HELM Arabic
-
10
HELM Arabic Enterprise
Simon Willison’s Weblog
-
1
Python 3.15.0 candidate 2 is here!
-
2
datasette-mcp 0.2
-
3
Quoting Tarn Adams
-
4
GeoJSON Map Viewer
-
5
Claude Fable 5.1 made me a really nice animated pelican
-
6
Quoting Rick Brewster
-
7
Claude's new system prompt really doesn't want to reproduce song lyrics
-
8
llm-gemini 0.34
-
9
Codex bundles LibreOffice
-
10
Quoting Andrew Digby
Silicon Republic
-
1
Blockdaemon adviser on the state of stablecoins
-
2
From blue to green: The former police officer turned sustainable engineer
-
3
HSE fined €645,000 for storing records in decrepit conditions
-
4
Viatel acquires UK network infrastructure specialist EDNX
-
5
Anthropic launches Claude Fable 5.1 and Mythos 5.1
-
6
Irish start-up Setanta plans €3m raise for its AI satellite technology
-
7
Uber cuts 3,300 jobs to simplify operations and boost productivity
-
8
33pc of employers pro remote hiring despite virtual interview concerns
-
9
Opera loses DMA ‘gatekeeper’ legal challenge over Microsoft Edge
-
10
Shein valued in IPO at $26bn – a quarter of its peak
Tech.eu
-
1
Motion lands $2M to expand humanoid robot deployments across Europe
-
2
NEURA Robotics acquires ADLATUS to bring Physical AI to autonomous cleaning
-
3
SoftBank invests $200M in Swiss robotics startup Gravis Robotics
-
4
Uber ups robotaxi offensive in Europe, with partnership expansion
-
5
NEURA Robotics acquires Bosch Rexroth’s ACTIVE Shuttle to expand Physical AI ecosystem
-
6
Kinematic Trees raises £585K to scale nature-inspired robotics software
-
7
Perceptual Robotics secures £4M+ to scale AI-powered wind inspections
-
8
UK robotics startup Humanoid hits $1.35B valuation with $152M Series A
-
9
Google Cloud and NVIDIA power microagi's embodied AI ambitions
-
10
Zalando joins Sereact's $116M Series B to accelerate AI-powered warehouse automation
Rest of World
-
1
Taiwan’s six-year hunt for China’s undercover chip labs
-
2
Meta accepts U.S. safety rules while pitching softer tools abroad
-
3
AI safety is designed in the West, and failing users everywhere
-
4
Why the global push to break free from Big Tech keeps falling short
-
5
India’s data center boom is leaving the people it displaces with nothing
-
6
I went to China to see a different AI future. It looked familiar
-
7
America’s immigration policy is driving away future AI leaders
-
8
The UAE is fighting AI hackers with AI of its own
-
9
A dumpling shop becomes a poster child of AI adoption in China
-
10
“It’s laughable”: Global AI experts challenge Zuckerberg’s “AI for everyone”
Together.ai
-
1
GLM-5.3 vs. GLM-5.3 Flash on DeepSWE: Cost, Coding, and Routing
-
2
GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
-
3
GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
-
4
DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing
-
5
A/B test models in production
-
6
DeepSeek V4 Pro 0813 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
-
7
DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding
-
8
Kimi K3: The Complete Developer Guide
-
9
Autoscaling endpoints for LLM inference
-
10
Together AI announces strategic partnership with Moonshot AI to natively serve Kimi models
Blog – neptune.ai
-
1
STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning [Paper Reflection]
-
2
How to Monitor, Diagnose, and Solve Gradient Issues in Foundation Models
-
3
SabiYarn: Advancing Low-Resource Languages With Multitask NLP Pre-Training [Paper Reflections]
-
4
Understanding Prompt Injection: Risks, Methods, and Defense Measures
-
5
Part 1: Instruction Fine-Tuning: Fundamentals, Architecture Modifications, and Loss Functions
-
6
A Researcher’s Guide to LLM Grounding
-
7
How to Optimize LLM Inference
-
8
Part 2: Instruction Fine-Tuning: Evaluation and Advanced Techniques for Efficient Training
-
9
Detecting and Fixing ‘Dead Neurons’ in Foundation Models
-
10
What are LLM Embeddings: All you Need to Know
Blog on EleutherAI Blog
-
1
What We Learned Trying to Catch AI Liars: An Aletheia's Quest Retrospective
-
2
A Dynamical Model of AI Governability
-
3
Rotary Embeddings: A Relative Revolution
-
4
Activation Function Ablation
-
5
Finetuning Models on Downstream Tasks
-
6
Evaluating Different Fewshot Description Prompts on GPT-3
-
7
On the Sizes of OpenAI API Models
-
8
Why Release a Large Language Model?
-
9
What A Long, Strange Trip It's Been: EleutherAI One Year Retrospective
-
10
Downstream Evaluations of Rotary Position Embeddings
Blog – PyImageSearch
-
1
Train YOLO26 on a Custom Dataset with YOLOE-26 Auto-Labeling
-
2
YOLO26 Open-Vocabulary Object Detection with YOLOE-26
-
3
Make a Chrome Extension to Digest Webpages with Manifest V3 and Groq API
-
4
Scaling, Optimizing, and Exporting Transformers with PyTorch Lightning
-
5
Training with PyTorch Lightning: Structured MLOps Development
-
6
Running Gemma 4 in the Browser with Transformers.js and WebGPU
-
7
Running Gemma 4 Locally: Ollama, llama.cpp, MLX, and More
-
8
Building Multimodal AI Applications with Gemma 4 and Transformers
-
9
Building a Multimodal Chatbot with Qwen3-VL Instruct and Thinking Models
-
10
Building an Intelligent Chatbot with Qwen3 Instruct and Thinking Models