<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Interpretability on Simone Sturniolo's Blog</title><link>https://stur86.github.io/s-plus-plus/tags/interpretability/</link><description>Recent content in Interpretability on Simone Sturniolo's Blog</description><generator>Hugo</generator><language>en-uk</language><lastBuildDate>Sat, 11 Jul 2026 06:35:24 +0100</lastBuildDate><atom:link href="https://stur86.github.io/s-plus-plus/tags/interpretability/index.xml" rel="self" type="application/rss+xml"/><item><title>Would You Kindly - Obedience vs truthfulness in local LLMs</title><link>https://stur86.github.io/s-plus-plus/posts/would-you-kindly/</link><pubDate>Sat, 11 Jul 2026 06:35:24 +0100</pubDate><guid>https://stur86.github.io/s-plus-plus/posts/would-you-kindly/</guid><description>&lt;p&gt;In Isaac Asimov&amp;rsquo;s &lt;em&gt;I, Robot&lt;/em&gt; classic story collection, the robots are subject to the famous &lt;a href="https://en.wikipedia.org/wiki/Three_Laws_of_Robotics"&gt;Three Laws of Robotics&lt;/a&gt;, which essentially set strict priorities for them that should always override all the lower ones (do not harm humans, obey orders, and protect yourself, in that order). In mathematical terms, such priorities would essentially be incomparable - &lt;em&gt;any&lt;/em&gt; level of concern about human welfare for example should instantly override any order. This is a problem in &lt;a href="https://en.wikipedia.org/wiki/Little_Lost_Robot"&gt;Little Lost Robot&lt;/a&gt;, for example, where a class of robots meant for use in a dangerous industrial environment causes more trouble than they&amp;rsquo;re worth by dragging away their own operators forcefully whenever they&amp;rsquo;re working around even entirely tolerable levels of background radiation. There&amp;rsquo;s no real way to formalize this kind of strict priority, and it would be impractical exactly for the reasons above - a smart enough robot/AI could conceive of virtually &lt;em&gt;any&lt;/em&gt; action it can take causing indirectly human harm, and therefore the only logical response would be paralysis. In practice, modern AI systems are built by optimising for one single loss function, however complex and stratified the training may be. This means that whatever priorities were set to it (for example, an LLM can be trained to reproduce text but also finetuned specifically to not respond rudely to its user), they will always be performing some kind of trade-off, operating on some Pareto frontier in which all different goals are weighted with respect to each other (and for example, it &lt;em&gt;is&lt;/em&gt; possible usually to &amp;ldquo;jailbreak&amp;rdquo; them into talking rudely to you if you create the conditions in which this seems the most appropriate way to reproduce the text accurately).&lt;/p&gt;</description></item></channel></rss>