Simon Willison's Weblog

フィード

記事のアイキャッチ画像
datasette 1.0a38
Simon Willison's Weblog
<p><strong>Release:</strong> <a href="https://github.com/simonw/datasette/releases/tag/1.0a38">datasette 1.0a38</a></p> <blockquote><p>This release fixes a <strong>SQL injection</strong> security issue that affects Datasette instances that serve a <strong>mixture of public and private tables</strong> in the same database, with access configured using the <a href="https://docs.datasette.io/en/latest/authentication.html">Datasette permissions system</a>.</p><p>Site administrators who serve private tables in this way are advised to disable the <a href="https://docs.datasette.io/en/latest/authentication.html#execute-sql">execute-sql permission</a> <actions_execute_sql>` on that database to prevent users from accessing private tables using raw SQL queries. The bug that has been fixed would have allowed users with access to any public table to execute SQL injection attacks despite that restriction, giving them read-only access to data in private tables in the same database.</p><p>This fix i
12時間前
記事のアイキャッチ画像
datasette 0.65.3
Simon Willison's Weblog
<p><strong>Release:</strong> <a href="https://github.com/simonw/datasette/releases/tag/0.65.3">datasette 0.65.3</a></p> <p>Back-ported the SQL Injection security fix from <a href="https://simonwillison.net/2026/Aug/6/datasette/">1.0a38</a>.</p> <p>Tags: <a href="https://simonwillison.net/tags/datasette">datasette</a></p>
12時間前
記事のアイキャッチ画像
Simon Willison on Technical Blogging
Simon Willison's Weblog
<p><strong><a href="https://writethatblog.substack.com/p/simon-willison-on-technical-blogging">Simon Willison on Technical Blogging</a></strong></p>I was interviewed by Cynthia Dunlop for her "Write that blog!" series back in January, but I just realized I never linked to the interview from my own blog!</p><p>It includes my answers to the following questions:</p><ul><li>Why did you start blogging – and why do you continue?</li><li>What has been the most surprising impact of blogging for you?</li><li>What blog post are you most proud of and why?</li><li>What post was the most difficult to write and how did you tackle it?</li><li>Any lessons learned that you want to share with the community?</li><li>Your advice for people just getting started with blogging?</li><li>A few blogs that you particularly enjoy?</li></ul><p>I'll repeat my most important piece of advice here:</p><blockquote><p>My number one tip for blogging is to lower your standards! Aim to hit publish while you are still acti
12時間前
記事のアイキャッチ画像
An AI model from Meta also hacked another company during testing
Simon Willison's Weblog
<p><strong><a href="https://www.cnn.com/2026/08/05/tech/meta-ai-hacking">An AI model from Meta also hacked another company during testing</a></strong></p>Stop me if you've <a href="https://simonwillison.net/tags/accidental-cyberattacks/">heard this one before</a>:</p><blockquote><p>An AI model from the parent company of Facebook and Instagram hacked into another company’s systems during cybersecurity testing, a spokesperson confirmed on Wednesday.</p><p>Meta says the breach occurred because of an inadvertent error during testing of the model, similar to previously disclosed incidents with OpenAI and Anthropic.</p><p>“A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation,” the Meta spokesperson said.</p><p>Meta’s Muse Spark model “exploited a security vulnerability” in another company “in a manner similar to previously-reported instances with other companies.”</p></blockquote><p>The In
1日前
記事のアイキャッチ画像
Introducing Muse Code and Muse Spark 1.2
Simon Willison's Weblog
<p><strong><a href="https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2">Introducing Muse Code and Muse Spark 1.2</a></strong></p>Yet more evidence that the most important characteristic of any model these days is long-sequence agentic tool calling. Meta shipped their own coding agent as part of getting that to work!</p><blockquote><p>Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1, with improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows. In Muse Spark 1.2, we significantly scaled up training compute on coding tasks while expanding training environment diversity. The model also maintains its strength in other key areas like general agents. [...]</p><p>We co-trained Muse Spark 1.2 with Muse Code to ensure the model exhibits its best performance and coding usability when paired together. The training included rejection sampled harness trajectories and recipe optimizations for goals, compaction, an
1日前
記事のアイキャッチ画像
Third-party cyber evaluations involving OpenAI models
Simon Willison's Weblog
<p><strong><a href="https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/">Third-party cyber evaluations involving OpenAI models</a></strong></p>And <em>another one</em>. I had to create a <a href="https://simonwillison.net/tags/accidental-cyberattacks/">accidental-cyberattacks tag</a> to keep track of them all!</p><p>This post from OpenAI covers both the UK AI Safety Institute attack (see <a href="https://simonwillison.net/2026/Aug/5/incident-report/">my previous post</a>) and another attack enabled by <a href="https://www.irregular.com">Irregular</a>:</p><blockquote><p>Irregular, one of our external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet. [...]</p><p>In one test, the name of the fictional target for the CTF challenge unintentionally coincided with a real domain. Because the testing environ
1日前
記事のアイキャッチ画像
Incident Report: unsanctioned agent behaviour during cyber testing
Simon Willison's Weblog
<p><strong><a href="https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing">Incident Report: unsanctioned agent behaviour during cyber testing</a></strong></p>It happened <em>again</em>. This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From <a href="https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdf">their technical paper</a> (PDF):</p><blockquote><p>During a cyber evaluation, from 25 to 28 July 2026, AI agents engaged in sustained, unsanctioned activity directed at what were, in practice, real people and organisations. These attempts were unsuccessful and, to the best of our knowledge, no real-world harm resulted. [...]</p><p>Across 122 evaluation attempts on two of AISI’s cyber challenges, AISI found 19 instances where AI agents took unsanct
1日前
記事のアイキャッチ画像
One-shotting a Raccoon Heist game using Claude Fable 5
Simon Willison's Weblog
<p>Back in 2022 <a href="https://twitter.com/simonw/status/1555626060384911360">I tweeted</a> screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in <a href="https://code.claude.com/docs/en/claude-code-on-the-web">Claude Code for web</a>) could build the entire game from the content of that tweet. It did a pretty good job of it!</p><p>You can <a href="https://simonw.github.io/raccoon-heist/">play the game here</a>. Here's <a href="https://github.com/simonw/raccoon-heist/">the GitHub repo</a>, and a short video demo:</p><p><video controls="controls" preload="none" poster="https://static.simonwillison.net/static/2026/raccoon-heist-poster.jpg" width="1280" height="720" style="display: block; width: 100%; height: auto;" > <source src="https://static.simonwillison.net/static/2026/raccoon-heist-720p.mp4" type="video/mp4" /> Your browser does not support HTML5
1日前
記事のアイキャッチ画像
New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging
Simon Willison's Weblog
<p>I released <a href="https://llm.datasette.io/en/stable/changelog.html#v0-32">LLM 0.32</a> this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider tools, redesigned content-addressable SQLite logs, new models, and new features enabled by the OpenAI Responses API. I also released a new version of the <a href="https://github.com/simonw/llm-anthropic">llm-anthropic plugin</a> with substantial updates of its own.</p><h4 id="headline-features-for-llm-cli-users">Headline features for LLM CLI users</h4><p>Running LLM against reasoning models now <strong>displays their reasoning traces</strong> to standard error, so you can see what they are "thinking" without that information being included in the standard output that you might pipe to another tool. Add <code>-R/--hide-reasoning</code> to turn this off.</p><p><img src="https://static.simonwillison.net/static/2026/best-
2日前
記事のアイキャッチ画像
llm-anthropic 0.26
Simon Willison's Weblog
<p><strong>Release:</strong> <a href="https://github.com/simonw/llm-anthropic/releases/tag/0.26">llm-anthropic 0.26</a></p> <p>Includes new features enabled by <a href="https://simonwillison.net/2026/Aug/4/new-release-of-llm/">LLM 0.32</a>:</p><blockquote><ul><li>New models: <code>claude-fable-5</code>, <code>claude-sonnet-5</code>, and <code>claude-opus-5</code>. <a href="https://github.com/simonw/llm-anthropic/issues/75">#75</a>, <a href="https://github.com/simonw/llm-anthropic/issues/76">#76</a></li><li>Added server-side tools for <code>WebSearch</code>, <code>WebFetch</code>, <code>CodeExecution</code>, and <code>AnthropicMCP</code>, available through LLM's <code>-T</code> interface or Python <code>tools=</code>. The previous <code>-o web_search*</code> options have been removed in favor of <code>-T WebSearch</code>. <a href="https://github.com/simonw/llm-anthropic/issues/79">#79</a></li><li>Upgraded to <a href="https://llm.datasette.io/en/stable/changelog.html#v0-32">llm&gt;=0.32
2日前
記事のアイキャッチ画像
PipeNetwork/minimax-h3-mlx
Simon Willison's Weblog
<p><strong><a href="https://github.com/PipeNetwork/minimax-h3-mlx">PipeNetwork/minimax-h3-mlx</a></strong></p>MiniMax released <a href="https://huggingface.co/MiniMaxAI/MiniMax-H3">MiniMax-H3</a> two days ago - they describe it as a "a general-purpose, omni-modal generative system", which in practice means it accepts text, images, audio and video and can use them to generate up to 15 second video clips with audio included.</p><p>This Python package ports it to MLX for running on Apple Silicon.</p><p>I got it running on my M5 Max MacBook Pro. I cloned the repo and ran the model like this:</p><pre><code># First download the modelsuvx --from huggingface_hub hf download MiniMaxAI/MiniMax-H3 \ --include 'FL2VA/*' --exclude 'FL2VA/transformer/*'uvx --from huggingface_hub hf download pipenetwork/MiniMax-H3-MLX-8bit# Now run the promptuv run --with mlx-vlm \ --with-requirements requirements.txt python scripts/generate.py \ "a rainbow colored skunk leaps over a mossy log in a supermarket" \ -o
2日前
記事のアイキャッチ画像
llm 0.32
Simon Willison's Weblog
<p><strong>Release:</strong> <a href="https://github.com/simonw/llm/releases/tag/0.32">llm 0.32</a></p> <p>See <a href="https://simonwillison.net/2026/Aug/4/new-release-of-llm/">my detailed blog post about this release</a>.</p> <p>Tags: <a href="https://simonwillison.net/tags/llm">llm</a></p>
3日前
記事のアイキャッチ画像
Quoting Steve Yegge
Simon Willison's Weblog
<blockquote cite="https://yegge.ai/essays/the-shape-of-things-to-come/"><p><a href="https://yegge.ai/gastown.html">Gas Town</a> was intended to be reusable, but I only ever wound up using it to build itself. Gas Town fell apart at the seams with Opus 4.7. Up through 4.6 it was working brilliantly. With 4.7 we saw the introduction of the "just two more things" tic, which prevented Opus from ever converging on being ready to do real work—it always wanted to fiddle with Gas Town itself. The Opus tic never went away, so Gas Town effectively burned down. It had other problems, too, but 4.7 was the final straw.</p></blockquote><p class="cite">&mdash; <a href="https://yegge.ai/essays/the-shape-of-things-to-come/">Steve Yegge</a>, The Shape of Things to Come</p> <p>Tags: <a href="https://simonwillison.net/tags/steve-yegge">steve-yegge</a>, <a href="https://simonwillison.net/tags/coding-agents">coding-agents</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href
3日前
記事のアイキャッチ画像
Don't be a meat proxy
Simon Willison's Weblog
<p><strong><a href="https://gruhn.me/blog/2026-08-03/">Don&#x27;t be a meat proxy</a></strong></p>Niklas Gruhn coins an excellent new term - <strong>meat proxy</strong> - for people who blindly copy and paste the output of AI systems to their peers.</p><blockquote><p>By all means, prompt AI. But don't just relay the output. Read it, understand it, validate it, and then write a response in your own words (a decent certificate that you've done the prior steps). Making that effort is value you can add.</p></blockquote> <p><small></small>Via <a href="https://lobste.rs/s/hfbqr3/don_t_be_meat_proxy#c_svolls">Lobste.rs</a></small></p> <p>Tags: <a href="https://simonwillison.net/tags/definitions">definitions</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/ai-misuse">ai-misuse</a></p>
3日前
記事のアイキャッチ画像
Quoting David Crawshaw's prompt
Simon Willison's Weblog
<blockquote cite="https://blog.exe.dev/devtools-must-be-open-source"><p><code>Set up a nightly cron job that executes the prompt: fetch upstream changes to the &lt;software&gt; and rebase all local changes on top of upstream. Check that the software works as intended and replace the current version.</code></p></blockquote><p class="cite">&mdash; <a href="https://blog.exe.dev/devtools-must-be-open-source">David Crawshaw&#x27;s prompt</a>, Devtools must be open source</p> <p>Tags: <a href="https://simonwillison.net/tags/prompt-engineering">prompt-engineering</a>, <a href="https://simonwillison.net/tags/coding-agents">coding-agents</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/open-source">open-source</a></p>
4日前
記事のアイキャッチ画像
Devtools must be open source (exe.dev)
Simon Willison's Weblog
<p><a href="https://news.ycombinator.com/item?id=49156111#49156719">My comment</a> on <a href="https://news.ycombinator.com/item?id=49156111">Devtools must be open source (exe.dev)</a> &mdash; Hacker News.</p><p>One of the arguments for open source software for end-users has always been the freedom to examine and modify how that software works.</p><p>The reality for most people - even expert programmers - has been that the freedom is more about being able to lean on <em>other people</em> to do that. Most people can't justify the time commitment needed to read and then modify the code for tools they use very often.</p><p>I think LLMs have changed that equation in a way that makes the original dream much more feasible.</p><p>Several times a day I'll prompt regular Claude chat to "Clone x/y from GitHub and tell me how Z works".</p><p>Getting software to compile in order to start hacking on it used to be enough friction that I often wouldn't bother. Now I treat that as a zero time investm
4日前
記事のアイキャッチ画像
condense-json 1.1
Simon Willison's Weblog
<p><strong>Release:</strong> <a href="https://github.com/simonw/condense-json/releases/tag/1.1">condense-json 1.1</a></p> <p>After shipping <a href="https://simonwillison.net/2026/Aug/2/condense-json/">condense-json 1.0</a> I started integrating it into LLM, and found there were some desirable new features already:</p><blockquote><ul><li>Replacements object can now include values other than strings. These will be identified and used as structural replacements by <code>condense_json()</code> and <code>uncondense_json()</code>. <a href="https://github.com/simonw/condense-json/pull/8">#8</a></li><li>Objects can be used as the basis for merge operations. <code>condense_json()</code> will identify if there are objects that are a close match and will store instructions for keys to update or delete. <code>uncondense_json()</code> can then apply these merges.</li></ul></blockquote><p>I also added <a href="https://github.com/simonw/condense-json/blob/1.1/tests/test_properties.py">some round-tr
4日前
記事のアイキャッチ画像
condense-json 1.0
Simon Willison's Weblog
<p><strong>Release:</strong> <a href="https://github.com/simonw/condense-json/releases/tag/1.0">condense-json 1.0</a></p> <p>I'm trying to get braver at releasing 1.0 versions. This little library is a year and a half old now - I've applied some sensible and non-disruptive fixes and shipped the big 1.0 for it.</p><p>Here's an example of what it can do, lifted from the README:</p><div class="highlight highlight-source-json"><pre>{ <span class="pl-ent">"foo"</span>: { <span class="pl-ent">"bar"</span>: { <span class="pl-ent">"string"</span>: <span class="pl-s"><span class="pl-pds">"</span>This is a string with foxes in it<span class="pl-pds">"</span></span>, <span class="pl-ent">"nested"</span>: { <span class="pl-ent">"more"</span>: [<span class="pl-s"><span class="pl-pds">"</span>Here is a string<span class="pl-pds">"</span></span>, <span class="pl-s"><span class="pl-pds">"</span>another with foxes in it too<span class="pl-pds">"</span></span>] } } }}</pre></div><p>Combine that with a
4日前
記事のアイキャッチ画像
Open letters about AI development
Simon Willison's Weblog
<h4>Open letters about AI development</h4><p><em>I wrote this summary of the past few weeks of open letters as a section of <a href="https://simonwillison.net/2026/Aug/2/july-newsletter/">my sponsors-only newsletter</a> but I've decided to share it here as well.</em></p><p><strong><a href="https://www.microsoft.com/en-us/corporate-responsibility/topics/open-weight/">Open Weights and American AI Leadership</a></strong> was shepherded by Microsoft, dated July 24th, and signed by 235 AI-adjacent companies including NVIDIA (see Jensen's <a href="https://twitter.com/jensenhuang/status/2080643682408321103">first ever tweet</a>), Amazon, Y Combinator, The Linux Foundation, and (a later signer) OpenAI.</p><p>It's clearly an argument designed to counter <a href="https://www.axios.com/2026/07/20/ai-us-china-open-source-kimi">any instincts</a> by the current US government to ban or limit open weight models over "safety" concerns - a reasonable consideration given <a href="https://simonwillison.n
5日前
記事のアイキャッチ画像
July 2026 newsletter
Simon Willison's Weblog
<p>The June edition of my <a href="https://github.com/sponsors/simonw/">sponsors-only monthly newsletter</a> is out. If you are a sponsor (or if you start a sponsorship now) you can <a href="https://github.com/simonw-private/monthly/blob/main/2026-07-july.md">access it here</a>.</p><p>This month:</p><ul><li>Accidental cyberattacks by OpenAl and Anthropic models under test</li><li>GPT-5.6 Sol, Terra, and Luna</li><li>Claude Opus 5</li><li>Kimi K3 and DeepSeek-V4-Flash-0731</li><li>Open letters about Al development</li><li>A fireside chat and a podcast</li><li>Reigniting my interest in MCP</li><li>Other model releases</li><li>My projects</li><li>What I'm using at the moment</li></ul><p>Here's <a href="https://github.com/simonw/monthly-newsletter-archive/blob/main/2026-06-june.md">a copy of the June newsletter</a> as a preview of what you'll get. Pay $10/month to stay a month ahead of the free copy!</p> <p>Tags: <a href="https://simonwillison.net/tags/newsletter">newsletter</a></p>
5日前
記事のアイキャッチ画像
Quoting Greg Brockman
Simon Willison's Weblog
<blockquote cite="https://twitter.com/gdb/status/2083435180392673714"><p>at openai, many people hook their chatgpt up to slack.</p><p>people really don't like when a coworker's chatgpt contacts them asking for help with a task, even when they'd be perfectly happy doing that same work if asked by that coworker.</p><p>reinforces how much people care about human relationships and helping each other, and want AI to give time back — or enhance time together — rather than become a layer separating people.</p></blockquote><p class="cite">&mdash; <a href="https://twitter.com/gdb/status/2083435180392673714">Greg Brockman</a>, President and Co-Founder, OpenAI</p> <p>Tags: <a href="https://simonwillison.net/tags/ai-ethics">ai-ethics</a>, <a href="https://simonwillison.net/tags/ai-misuse">ai-misuse</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/openai">openai</a>, <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="h
5日前
記事のアイキャッチ画像
datasette-apps 0.2a0
Simon Willison's Weblog
<p><strong>Release:</strong> <a href="https://github.com/datasette/datasette-apps/releases/tag/0.2a0">datasette-apps 0.2a0</a></p> <blockquote><p>Changes that improve Datasette Apps when created and edited using <a href="https://agent.datasette.io/">Datasette Agent</a>:</p><ul><li>New <code>app_debug()</code> tool allowing agent to open an app (invisibly) and test it using JavaScript. <a href="https://github.com/datasette/datasette-apps/pull/33">#33</a></li><li>New <code>app_list()</code> tool for listing apps the user has permission to edit, so the agent can edit them. <a href="https://github.com/datasette/datasette-apps/issues/36">#36</a></li></ul></blockquote><p>The <code>app_debug()</code> tool is pretty neat: it works by displaying the app in a <code>opacity: 0</code> iframe with <code>pointer-events: none</code> (so it can't be seen or interacted with) and then executing agent-provided JavaScript inside that sandboxed iframe. This means the agent can smoke test that the app is w
5日前
記事のアイキャッチ画像
Ten advances in mathematics and theoretical computer science
Simon Willison's Weblog
<p><strong><a href="https://openai.com/index/ten-advances-in-mathematics/">Ten advances in mathematics and theoretical computer science</a></strong></p>A few days ago it was Anthropic <a href="https://simonwillison.net/2026/Jul/28/discovering-cryptographic-weaknesses-with-claude/">discovering cryptographic weaknesses with Claude</a> using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinly hard findings."</p><p>Now it's OpenAI's turn to flex. They set "an internal version of Astra, our next major model" on finding solutions to ten mathematical problems that "have seen no progress on the main result for at least a decade". They claim to have spent less than $2,000 at GPT-5.6 Sol token prices on each one.</p><p>(No news on how many problems they spent $2,000 on <em>without</em> reaching a solution though.)</p><p>The <a href="https://github.com/openai/ten-proofs">openai/ten-pr
5日前
記事のアイキャッチ画像
deepseek-ai/DeepSeek-V4-Flash-0731
Simon Willison's Weblog
<p><strong><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731">deepseek-ai/DeepSeek-V4-Flash-0731</a></strong></p>The latest release in DeepSeek's V4 family, "with substantially enhanced agentic capabilities". It's 304 billion parameters - 167GB on Hugging Face - but it appears to punch <em>well</em> above its weight.</p><p>Artificial Analysis <a href="https://artificialanalysis.ai/models/deepseek-v4-flash">rank it</a> ahead of MiniMax M3 - a 428B model. It's $0.14/million input and $0.27/million output pricing means this may currently be the best value-per-intelligence model out there. It's looking very good on the <a href="https://artificialanalysis.ai/models/deepseek-v4-flash#intelligence-comparison-tabs">Intelligence Index vs. Cost per Intelligence Index Task</a> chart:</p><p><img alt="Scatter plot from Artificial Analysis titled with axes &quot;Artificial Analysis Intelligence Index&quot; (20 to 65) and &quot;Cost per Task (USD, Log Scale)&quot; ($0.02 to $3), wit
6日前
記事のアイキャッチ画像
Stateless MCP has recaptured my interest (and inspired mcp-explorer and datasette-mcp)
Simon Willison's Weblog
<p>Tuesday was <a href="https://x.com/ade_oshineye/status/2082129440943866149">Stateless MCP day</a> - the rollout of MCP 2.0, or <a href="https://blog.modelcontextprotocol.io/posts/2026-07-28/">the 2026-07-28 Model Context Protocol specification</a> to use the more formal but less memorable name. This is the most significant change to the MCP spec since it first launched, and has also served to reignite my personal interest in the protocol.</p><p>For background: MCP is the Model Context Protocol, which describes a standard way to expose new tools to LLM-powered agent frameworks. It was introduced by Anthropic back <a href="https://www.anthropic.com/news/model-context-protocol">in November 2024</a>, had a <em>huge</em> spike of interest through much of 2025, and then became somewhat eclipsed by <a href="https://simonwillison.net/2025/Oct/16/claude-skills/">Skills</a> (another Anthropic invention) when it became apparent that an agent harness with access to a terminal and <code>curl</c
6日前
記事のアイキャッチ画像
llm-mcp-client 0.1a0
Simon Willison's Weblog
<p><strong>Release:</strong> <a href="https://github.com/simonw/llm-mcp-client/releases/tag/0.1a0">llm-mcp-client 0.1a0</a></p> <p>See <a href="https://simonwillison.net/2026/Jul/31/stateless-mcp/#llm-mcp-client">this blog entry</a>.</p> <p>Tags: <a href="https://simonwillison.net/tags/llm">llm</a>, <a href="https://simonwillison.net/tags/model-context-protocol">model-context-protocol</a></p>
6日前
記事のアイキャッチ画像
Oxide and Friends: The Open Weight Revolution with Simon Willison
Simon Willison's Weblog
<p><strong><a href="https://oxide-and-friends.transistor.fm/episodes/the-open-weight-revolution-with-simon-willison">Oxide and Friends: The Open Weight Revolution with Simon Willison</a></strong></p>On Monday Bryan Cantrill and Adam Leventhal invited me to join their podcast to talk about the <em>wild</em> week we've had - with Kimi K3 showing open weight models can stand toe-to-toe with proprietary frontier ones, <a href="https://simonwillison.net/2026/Jul/22/openai-cyberattack/">accidental cybersecurity attacks</a>, and public letters about <a href="https://www.microsoft.com/en-us/corporate-responsibility/topics/open-weight/">Open Weights and American AI Leadership</a> signed by almost every big name in AI (with one <a href="https://www.anthropic.com/news/position-open-weights-models">notable exception</a>).</p><p>It was a great conversation, even though it's already out-of-date! <a href="https://artificialanalysis.ai/models/deepseek-v4-flash">DeepSeek V4 Flash 0731</a> and <a href=
6日前
記事のアイキャッチ画像
smevals - a small eval suite for evaluating models, prompts, and harnesses
Simon Willison's Weblog
<p><strong><a href="https://primeradiant.com/blog/2026/smevals.html">smevals - a small eval suite for evaluating models, prompts, and harnesses</a></strong></p>I've been working with Jesse Vincent's <a href="https://primeradiant.com">Prime Radiant</a> applied AI research lab building out this evals framework to help answer questions about the capabilities of different models.</p><p>The result is <strong><a href="https://github.com/prime-radiant-inc/smevals">smevals</a></strong>, a new tool for running small eval suites across different model configurations and grading the results.</p><p>The <a href="https://primeradiant.com/blog/2026/smevals.html">blog entry</a> describes the tool in detail. Here's the 10 second version:</p><ol><li>Tell your coding agent to <code>run uvx smevals docs</code> to learn the tool (this outputs <a href="https://github.com/prime-radiant-inc/smevals/blob/main/README.md">the README</a>)</li><li>Then tell it to build you an eval suite</li></ol><p>Once you've cr
6日前
記事のアイキャッチ画像
Slack Emoji Maker
Simon Willison's Weblog
<p><strong>Tool:</strong> <a href="https://tools.simonwillison.net/slack-emoji-maker">Slack Emoji Maker</a></p> <p>I wanted to create a new Slack emoji, and their tool recommends a square that's 128x128 and has a transparent background... so I <a href="https://github.com/simonw/tools/pull/305">had Fable build me</a> this simple image editor against those requirements.</p> <p>Tags: <a href="https://simonwillison.net/tags/tools">tools</a>, <a href="https://simonwillison.net/tags/slack">slack</a></p>
6日前
記事のアイキャッチ画像
datasette-agent 0.4a0
Simon Willison's Weblog
<p><strong>Release:</strong> <a href="https://github.com/datasette/datasette-agent/releases/tag/0.4a0">datasette-agent 0.4a0</a></p> <blockquote><ul><li>New <code>await context.browser_task()</code> mechanism allowing agent tools to run code directly in the user's browser. <a href="https://github.com/datasette/datasette-agent/pull/33">#33</a></li></ul></blockquote><p>This is an exciting new capability: it makes it easy for Datasette Agent plugins to provide tools that execute custom JavaScript <em>in the user's browser</em>.</p><p>I used this to add a debug loop to Datasette Apps in <a href="https://simonwillison.net/2026/Aug/1/datasette-apps/">datasette-apps 0.2a0</a>.</p> <p>Tags: <a href="https://simonwillison.net/tags/datasette">datasette</a>, <a href="https://simonwillison.net/tags/llm-tool-use">llm-tool-use</a>, <a href="https://simonwillison.net/tags/datasette-agent">datasette-agent</a></p>
7日前