<?xml version="1.0" encoding="utf-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Posts · Mesh Refinement</title><link>https://meshrefine.com/en/</link><description>Posts · Mesh Refinement</description><atom:link href="https://meshrefine.com/en/index.xml" rel="self" type="application/rss+xml"/><lastBuildDate>Mon, 31 Aug 2026 20:21:50 +0200</lastBuildDate><item><title>Chronicles. Aug. 24 - Aug. 30 2026</title><link>https://meshrefine.com/en/posts/chronicles-aug-24-aug-30-2026/</link><description>&lt;p&gt;It&amp;rsquo;s Monday, and so it&amp;rsquo;s time for the next issue of Chronicles. Last week was very significant because it represented a tectonic shift from using expensive, efficient models to cheap and still very efficient models. In other words, AI is getting democratized.&lt;/p&gt;
&lt;p&gt;Let&amp;rsquo;s start with the hardware. This week we&amp;rsquo;ve seen interesting news from OpenAI, which &lt;a href="https://openai.com/index/jalapeno-first-results"&gt;published the first benchmarks&lt;/a&gt; for their Jalapeño chip, a custom inference chip designed with Broadcom. It&amp;rsquo;s pretty impressive. They published &lt;a href="https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia"&gt;over 700 tokens per second per user on DeepSeek R1 at concurrency of one, and about 1400 on Kimi K2.5 and GPT-OSS&lt;/a&gt;. The variety here matters because the chip has been benchmarked not only on OpenAI&amp;rsquo;s own models but on open-weight models from other vendors as well. The chip is critical in the ongoing price war with the Chinese providers. Remember that last week we had data on &lt;a href="https://community.openai.com/t/20-price-reduction-for-gpt-5-6-sol-api-codex-credits-and-chatgpt-work/1391726"&gt;cutting prices on their Sol model&lt;/a&gt; by 20% on input and 33% on output, which is a promotional window that expires on 21 November, and a week before that they &lt;a href="https://arstechnica.com/ai/2026/08/openai-and-anthropic-in-price-war-as-chinese-ai-rivals-gain-ground/"&gt;cut the price of Luna by a whopping 80%&lt;/a&gt;.&lt;/p&gt;</description><pubDate>Mon, 31 Aug 2026 20:21:50 +0200</pubDate><guid>https://meshrefine.com/en/posts/chronicles-aug-24-aug-30-2026/</guid></item><item><title>Chronicles. Aug. 17 - Aug. 23 2026</title><link>https://meshrefine.com/en/posts/chronicles-aug-17-aug-23-2026/</link><description>&lt;p&gt;Last week was all about the harness. Gone are the times when the models were all things-in-themselves. Now, as with early Homo erectus, the tools are the factor of survival.
If we look at the types of harness-related posts and announcements, we will see three trends.&lt;/p&gt;
&lt;h2 id="04602056b054f8f419959cfb26ab2e8f-harness-products"&gt;Harness products&lt;/h2&gt;
&lt;p&gt;Several companies have released either agents or infra for agents at once.&lt;/p&gt;
&lt;p&gt;OpenAI has &lt;a href="https://developers.openai.com/blog/codex-as-a-platform"&gt;released the execution framework&lt;/a&gt; that their Codex (the CLI one) is based on. Unsurprisingly called Harness, it is their answer to the &lt;a href="https://code.claude.com/docs/en/agent-sdk/overview"&gt;Claude Agent SDK&lt;/a&gt; and the &lt;a href="https://docs.github.com/en/copilot/how-tos/copilot-sdk"&gt;GitHub Copilot SDK&lt;/a&gt;. The Harness provides the execution loop, memory, tools, and other necessities of agentic life. You can use it in three different ways. First, you can just use &lt;code&gt;codex exec&lt;/code&gt; to run non-interactive jobs. Second, you can use the Codex SDK for building workflows. And last, you can use app-server to build apps that require conversation handling.
One point worthy of attention is that they claim that the Harness improved the performance of the GPT-5.6 Sol model from 13.3% to 38.3% on ARC-AGI-3. It&amp;rsquo;s unclear what harness (not capitalized) was used as a baseline, though.&lt;/p&gt;</description><pubDate>Mon, 24 Aug 2026 20:03:30 +0200</pubDate><guid>https://meshrefine.com/en/posts/chronicles-aug-17-aug-23-2026/</guid></item><item><title>Chronicles. Aug. 09 - Aug. 16 2026</title><link>https://meshrefine.com/en/posts/chronicles-aug-09-aug-16-2026/</link><description>&lt;p&gt;Over the last week, we&amp;rsquo;ve seen some announcements from both the highest-end and lowest-end sides of open-weight models. At the same time, labs are starting to think that maaaybe, just maybe, we move too fast and we need to stop and think. So, in their free time, they are starting price wars.&lt;/p&gt;
&lt;h2 id="d8260bc688ef019c5f40c851636b0844-models-hi"&gt;Models, hi&lt;/h2&gt;
&lt;p&gt;Last week we saw a bunch of releases of very large open-weight models. Z.ai shipped &lt;a href="https://z.ai/blog/glm-5.3"&gt;GLM-5.3&lt;/a&gt;, which is relatively small, just 743B, DeepSeek published their 1.6T &lt;a href="https://api-docs.deepseek.com/news/news260813/"&gt;V4-Pro&lt;/a&gt; weights under the MIT license, and Alibaba surprised with a whopping 2.4-trillion-parameter &lt;a href="https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B"&gt;Qwen3.8&lt;/a&gt;. Although they are open, doing anything meaningful with models of this size requires hardware not a lot of individuals have, which defeats their openness a little. On the other hand, they &lt;em&gt;open&lt;/em&gt; possibilities for 3rd-party hosting and incentivize the price wars I will talk about below.&lt;/p&gt;</description><pubDate>Sun, 16 Aug 2026 09:31:26 +0200</pubDate><guid>https://meshrefine.com/en/posts/chronicles-aug-09-aug-16-2026/</guid></item><item><title>Chronicles. Aug. 01 - Aug. 08 2026</title><link>https://meshrefine.com/en/posts/chronicles-aug-01-aug-08-2026/</link><description>&lt;p&gt;This article opens a series (I expect it to become one) of posts in which I review the AI-related events of the past week and try to figure out how they fit into the bigger picture.&lt;/p&gt;
&lt;p&gt;The main highlights of the past week are:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;We continue to see real-world security breaches caused by AI system evaluations, and it looks like another race.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Major regulations came into force.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Data centres face increasing opposition from local communities. Companies seek workarounds.&lt;/p&gt;</description><pubDate>Sun, 09 Aug 2026 07:47:17 +0200</pubDate><guid>https://meshrefine.com/en/posts/chronicles-aug-01-aug-08-2026/</guid></item><item><title>Clean your docs first</title><link>https://meshrefine.com/en/microposts/2025-08-29-documents-sanitization/</link><description>&lt;p&gt;Lately, there have been numerous alerts about security vulnerabilities connected to indirect prompt injection attacks. The main conclusion is: LLMs are gullible, nobody knows how to make them completely robust. Therefore, no AI system is safe.&lt;/p&gt;
&lt;p&gt;It is, of course, completely true, and the attention those attacks attract is most definitely welcome, but the main question is left unanswered: what is to be done about it?&lt;/p&gt;
&lt;p&gt;Purists would push for banning all external inputs that could lead to prompt injection, but, frankly, I wouldn&amp;rsquo;t be so rigid. A lot of genuinely useful applications do rely on external input, and so, I am afraid, we have to retreat to the last resort: engineering discipline.&lt;/p&gt;
&lt;p&gt;There are various schemes that use a clever interaction of different models to minimize their exposure to attacks (see, for example, &lt;a href="https://arxiv.org/abs/2503.18813"&gt;the CaMeL paper&lt;/a&gt;). I believe, though, that we don&amp;rsquo;t pay enough attention to simple input sanitization.&lt;/p&gt;
&lt;p&gt;A lot of such attacks rely on text hidden from the human, but visible to the machine. Detecting and removing such text is relatively trivial (in a programmatic sense), although it will require working not on the text but on the container level. This technique (called &lt;a href="https://en.wikipedia.org/wiki/Content_Disarm_%26_Reconstruction"&gt;Content Disarm &amp;amp; Reconstruction&lt;/a&gt;, by the way) has been around for quite some time. Honestly, I am surprised that it is not implemented everywhere. It wouldn&amp;rsquo;t prevent all attacks, but it would make an unsophisticated attacker&amp;rsquo;s life harder.&lt;/p&gt;
&lt;p&gt;And to those who insist that 99% in security is a failing score, I would like to remind them of a fundamental security principle: no system is 100% secure. The role of security is to make an attack more costly than the potential benefit gained from it (see &lt;a href="https://en.wikipedia.org/wiki/Gordon%E2%80%93Loeb_model"&gt;the Gordon–Loeb model&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;With this paradigm in mind, we &lt;em&gt;can&lt;/em&gt; build reliable and secure AI systems, even with insecure individual components.&lt;/p&gt;</description><pubDate>Fri, 29 Aug 2025 12:23:20 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-08-29-documents-sanitization/</guid></item><item><title>2025-08-08 09:45</title><link>https://meshrefine.com/en/microposts/2025-08-08-about-gpt-5/</link><description>&lt;p&gt;Some of the things you need to know about the latest GPT-5 release that evangelists don&amp;rsquo;t talk about:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;GPT-5 is not one model&lt;/strong&gt;. What they call GPT-5 outside of API context is a router that sends your request to a model that it thinks would work most efficiently on it. You need to look at the OpenAI&amp;rsquo;s promise to provide access to everyone in this light. They provide access to the router, and you don&amp;rsquo;t know the specific configuration it applies to you and whether you will actually be able to test the most powerful model. That would undoubtedly cause completely different experiences for different users.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;GPT-5 is not a PhD&lt;/strong&gt;. It is a pretty capable model (at least, one of them) that excels at &lt;em&gt;some&lt;/em&gt; tasks. You can expect improvements in:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;coding capabilities (they are impressive according to some vibe tests, but according to the SWE-bench Verified benchmark, it has just a minor lead over Claude Opus 4.1);&lt;/li&gt;
&lt;li&gt;tool calling capabilities, which are the most important for agentic workloads;&lt;/li&gt;
&lt;li&gt;other tasks where OpenAI had access to immediate feedback.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;While those are important improvements for a lot of areas, they don&amp;rsquo;t make the model PhD-level. The &lt;a href="https://bren.blog/gpt-5-demo-mistake-about-bernoulli-effect"&gt;hallucination about the airfoil&lt;/a&gt; during the demo perfectly demonstrates that it still internalizes &lt;em&gt;the most common&lt;/em&gt; belief, not the most current. It is a very hard problem to solve and actually a major roadblock on the way to AGI.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Diminished hallucinations are a double-edged sword&lt;/strong&gt;. On the one hand, the less a model hallucinates, the better, as you can trust it more. On the other hand, the more you trust the model, the more likely you are to miss actual hallucinations. In the real world, the model that never hallucinates is the best. The model that hallucinates in 0.01% of cases can be more dangerous than one that hallucinates in 10% of cases.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;My personal impression so far is that it still has the same issue as the previous models from OpenAI, namely that it is really superficial without careful prompting. It provides you with the most shallow analysis it can get away with and hides this fact by using the very well-structured responses.&lt;/p&gt;</description><pubDate>Fri, 08 Aug 2025 09:45:22 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-08-08-about-gpt-5/</guid></item><item><title>2025-07-31 11:44</title><link>https://meshrefine.com/en/microposts/2025-07-31-google-opal/</link><description>&lt;p&gt;&lt;a href="https://developers.googleblog.com/en/introducing-opal/"&gt;Google Opal&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;
&lt;picture&gt;
&lt;source srcset="https://meshrefine.com/en/microposts/2025-07-31-google-opal/google-opal_hu_df8c6253ecad9107.webp" type="image/webp" /&gt;
&lt;img class="img-fluid" src="https://meshrefine.com/en/microposts/2025-07-31-google-opal/google-opal.f5b0b4e4310d106f8f86ef27ed3e0029.jpg" alt="Google Opal" loading="lazy" height="723" width="1374" /&gt;
&lt;/picture&gt;
&lt;/p&gt;
&lt;p&gt;Google has started a public preview for its new tool for graphical creation of multi-step AI workflows. While not a tool for production use, it is a great helper for building personal tools (and everybody should build personal tools, really, it is the biggest differentiator now).&lt;/p&gt;
&lt;p&gt;It works on the Gemini platform with Gemini models (well, of course), and since it is a preview tool, it is free for now. In the future it will most likely use the Gemini Plan if the user has it. I am not sure about custom API keys, as this looks like a general public tool, but we&amp;rsquo;ll see.&lt;/p&gt;
&lt;p&gt;Right now it is available only in the US.&lt;/p&gt;</description><pubDate>Thu, 31 Jul 2025 11:44:43 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-31-google-opal/</guid></item><item><title>2025-07-31 10:19</title><link>https://meshrefine.com/en/microposts/2025-07-31-horizon-alpha-model/</link><description>&lt;p&gt;&lt;a href="https://x.com/OpenRouterAI/status/1950713168193282078"&gt;X is buzzing&lt;/a&gt; with a new Horizon Alpha model that is beating all previous models on various vibe tests (read: unicorns on bicycles and so on) singlehandedly. This model has 256k context window, which is a solid, albeit not the most impressive, number.&lt;/p&gt;
&lt;p&gt;Most probably it is a new OpenAI model (GPT-5?), as they already did the same trick before. You can try it on &lt;a href="https://openrouter.ai/openrouter/horizon-alpha"&gt;openrouter.ai&lt;/a&gt; completely for free, but remember not to provide it with any private data, as it is collected and used for model improvement.&lt;/p&gt;</description><pubDate>Thu, 31 Jul 2025 10:19:46 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-31-horizon-alpha-model/</guid></item><item><title>2025-07-28 11:01</title><link>https://meshrefine.com/en/microposts/2025-07-28-claude-code-subagent/</link><description>&lt;p&gt;&lt;a href="https://docs.anthropic.com/en/docs/claude-code/sub-agents"&gt;Claude Code Sub Agent&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Anthropic has added the ability to create and use specialized sub-agents in Claude Code. These sub-agents use a separate context window, which allows you to run separate tasks without polluting the main context, limiting the context rot effect. You can run the created sub-agents manually or let Claude Code decide when to use them.&lt;/p&gt;
&lt;p&gt;What can be a good sub-agent? Anything that needs to be an expert in its area, doesn&amp;rsquo;t need to share the context with the main agent, can have dedicated tools, and can be designed to run self-sufficient tasks. To give you a taste of what can be made into a sub-agent, here are a couple of examples:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Git sub-agent to which you can offload various git-related operations. As it can use just the local git state, and doesn&amp;rsquo;t need to have access to the global context, it is a good choice for a tool with its own context window.&lt;/li&gt;
&lt;li&gt;A book-writing sub-agent (I know, but I create book-like documents in a specific style for self-education). You provide it with a high-level plan, a set of materials, and one or two sections as examples, and let it work on a section separately from other sub-agents.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The second example would really benefit from the ability to run several sub-agents in parallel, but it looks like I&amp;rsquo;m asking too much.&lt;/p&gt;</description><pubDate>Mon, 28 Jul 2025 11:01:28 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-28-claude-code-subagent/</guid></item><item><title>2025-07-21 14:52</title><link>https://meshrefine.com/en/microposts/2025-07-21-quot1/</link><description>&lt;p&gt;&lt;a href="https://x.com/random_walker/status/1946180439045018046"&gt;Quoting Arvind Narayanan&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;If we compared AI capabilities against humans with no access to tools, such as the internet, we would probably find that AI already outperformed humans at many or most cognitive tasks we perform at work. But of course this is not a helpful comparison and doesn’t tell us much about AI’s economic impacts. &lt;strong&gt;We are nothing without our tools&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;</description><pubDate>Mon, 21 Jul 2025 14:52:06 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-21-quot1/</guid></item><item><title>2025-07-20 18:40</title><link>https://meshrefine.com/en/microposts/2025-07-20-twelvelabs-models-in-amazon-bedrock/</link><description>&lt;p&gt;&lt;a href="https://aws.amazon.com/blogs/aws/twelvelabs-video-understanding-models-are-now-available-in-amazon-bedrock/"&gt;TwelveLabs video understanding models are now available in Amazon Bedrock&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;AWS adds native video embeddings and video understanding models to Amazon Bedrock. It opens a lot of potential use cases for which I previously reached for Gemini models. One example of such a case is an educational system that watches how the learner performs the task and provides feedback based on the educational materials.&lt;/p&gt;
&lt;p&gt;Bedrock had workflows to do video understanding, but it was exactly that: workflows, not native models. You can imagine what they looked like&amp;mdash;take a video, split to frames, feed frames to VLM, try to maintain temporal consistency, despair, come to terms with the system&amp;rsquo;s performance, and go on vacation.&lt;/p&gt;
&lt;p&gt;Now, however, there are not one, but two different native video models:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;TwelveLabs Marengo, for creating video embeddings;&lt;/li&gt;
&lt;li&gt;TwelveLabs Pegasus, for video-based text generation.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Pricing of the models depends on whether your video has an audio track or not, but you should expect $2.5-$3/hour of video for Marengo and $1.8/hour for Pegasus.&lt;/p&gt;</description><pubDate>Sun, 20 Jul 2025 18:40:39 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-20-twelvelabs-models-in-amazon-bedrock/</guid></item><item><title>2025-07-20 10:26</title><link>https://meshrefine.com/en/microposts/2025-07-20-context-rot-self-reinforcing-structure/</link><description>&lt;p&gt;One form of &lt;a href="https://simonwillison.net/2025/Jun/18/context-rot/"&gt;context rot&lt;/a&gt; is what I call self-reinforced structure. When you accept a long-form model answer, you signal that this structure is acceptable, and so it tries to generate subsequent responses in a similar way. It can be destructive for any long-form creative work.&lt;/p&gt;
&lt;p&gt;The only real defense is ensuring that the history the model receives doesn&amp;rsquo;t contain such replies. So it should be either prevented early or later fixed by providing a summary of the previous conversation instead of the actual history.&lt;/p&gt;</description><pubDate>Sun, 20 Jul 2025 10:26:39 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-20-context-rot-self-reinforcing-structure/</guid></item><item><title>2025-07-17 22:39</title><link>https://meshrefine.com/en/microposts/2025-07-17-chatgpt-agent/</link><description>&lt;p&gt;&lt;a href="https://openai.com/index/introducing-chatgpt-agent/"&gt;Introducing ChatGPT agent: bridging research and action&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;OpenAI released an agent that can control your own computer. It uses its advanced reasoning capabilities to plan and solve tasks in applications like Excel and PowerPoint.&lt;/p&gt;
&lt;p&gt;While Sam Altman &amp;ldquo;feels the agi&amp;rdquo; looking at how the system works, I find it incredibly clunky. Instead of concentrating on providing the models with native tools (MCP is a good step forward, although not without its problems), they try to emulate hands and eyes for them, so models can do the same things we do, but slowly and awkwardly.&lt;/p&gt;
&lt;p&gt;So I would consider this type of agent a temporary workaround until we develop better machine-to-machine communication mechanisms. After this, it will be used to serve an increasingly long tail of legacy systems that will not have such machine-usable interfaces.&lt;/p&gt;
&lt;p&gt;P.S. Gemini mentioned a point of view I didn&amp;rsquo;t consider, namely that such systems can collect data necessary to train better embodied intelligence, meaning one that can act in the real world. It&amp;rsquo;s a perfectly valid point that shouldn&amp;rsquo;t be left without attention.&lt;/p&gt;</description><pubDate>Thu, 17 Jul 2025 22:39:57 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-17-chatgpt-agent/</guid></item><item><title>2025-07-17 20:40</title><link>https://meshrefine.com/en/microposts/2025-07-17-stanfordhai2025/</link><description>&lt;p&gt;&lt;a href="https://hai.stanford.edu/ai-index/2025-ai-index-report"&gt;Stanford&amp;rsquo;s 2025 AI Index Report&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Stanford published its annual report. It&amp;rsquo;s pretty important, because it separates speculation from pure numbers. Along with some obvious things (AI is getting better, cheaper, widespread, duh), there are some very interesting facts:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;While almost every organization is using AI now (78% in 2024, although no doubt, for most of them it boils down to using chatbots to compose emails), the actual results are somewhat modest. The productivity increase is on the scale of 10% (to be honest, such an increase in one year is kinda unprecedented), but the increase in revenue for most industries is just about 5%. Why? Because as with any general purpose technology, realization of full benefit would require complete rebuilding the organizational structures and processes. The problem is that no one knows how these new processes would look like, and we will have to learn from our own mistakes.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Maybe old news, but AI provides more leverage to less experienced employees. The great equalizer of modern times. Again, that means that we need to reformulate our approach to team staffing. I would only add that it can help only if you have some remote understanding of what you&amp;rsquo;re doing, so those who apply for entry positions, do your homework well.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The number of AI-related incidents continues to rise. We see a twofold increase in 2024 vs 2023, and this is before frantic adoption of Agents and MCPs we see in 2025. So, we need to brace ourselves and be ready for more and more &lt;a href="https://simonwillison.net/2025/Jul/6/supabase-mcp-lethal-trifecta/"&gt;data leaks&lt;/a&gt; and integrity breaches.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The report contains a lot more nuggets, but it&amp;rsquo;s almost 500 pages long, so I would really recommend to use AI to extract what you fancy.&lt;/p&gt;</description><pubDate>Thu, 17 Jul 2025 20:40:46 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-17-stanfordhai2025/</guid></item><item><title>2025-07-17 13:36</title><link>https://meshrefine.com/en/microposts/2025-07-17-voxtral/</link><description>&lt;p&gt;&lt;a href="https://mistral.ai/news/voxtral"&gt;Voxtral&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Mistral introduces Voxtral, a family of open-source speech recognition and understanding models. It&amp;rsquo;s about time. We haven&amp;rsquo;t seen a comparable open-source model since OpenAI&amp;rsquo;s Whisper, and that was quite a while ago.&lt;/p&gt;
&lt;p&gt;The models are provided in 3B and 24B sizes and outperform Whisper on most benchmarks. However, they require more powerful hardware, as the largest Whisper variant is just 1.5B. This is a direct consequence of it also being a regular language model. Another consequence is that controlling them in a pure transcription setting would be harder.&lt;/p&gt;
&lt;p&gt;The models are available on &lt;a href="https://huggingface.co/mistralai/"&gt;Hugging Face&lt;/a&gt; as well as through the Mistral API and their &lt;a href="https://chat.mistral.ai/"&gt;LeChat&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;What they also currently lack is diarization (speaker recognition) support. It&amp;rsquo;s on the roadmap, but in the meantime, we still have to use somewhat clunky &lt;a href="https://github.com/pyannote/pyannote-audio"&gt;pyannote-audio&lt;/a&gt; for this purpose.&lt;/p&gt;</description><pubDate>Thu, 17 Jul 2025 13:36:29 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-17-voxtral/</guid></item><item><title>2025-07-17 12:16</title><link>https://meshrefine.com/en/microposts/2025-07-17-ai-prisoners-dilemma/</link><description>&lt;p&gt;&lt;a href="https://arxiv.org/abs/2507.02618"&gt;Strategic Intelligence in Large Language Models: Evidence from evolutionary Game Theory&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The paper shows that different models behave completely differently when placed in game theory settings. What that means is that testing and evals are playing an increasingly critical role in developing agentic systems, as updating or changing the underlying model will lead to unpredictable changes in an agent&amp;rsquo;s behaviour.&lt;/p&gt;</description><pubDate>Thu, 17 Jul 2025 12:16:48 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-17-ai-prisoners-dilemma/</guid></item><item><title>2025-07-15 11:43</title><link>https://meshrefine.com/en/microposts/2025-07-15-aws-introduces-kiro/</link><description>&lt;p&gt;&lt;a href="https://kiro.dev/blog/introducing-kiro/"&gt;Introducing Kiro&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;AWS jumps into the agentic IDEs bandwagon with Kiro. To separate itself from vibe-coding approach, which accumulated a considerable amount of ill repute, they emphasize the &amp;ldquo;spec-driven development&amp;rdquo; method. That means that the agent first helps the user to create a full requirements document for the feature, then it analyzes the existing code base, and only after that it starts implementing.&lt;/p&gt;
&lt;p&gt;This approach definitely makes sense, and it&amp;rsquo;s a step forward from blindly running into the fray that is vibe-coding. The fact that those specs are updating along with the code changes makes them even more valuable, minimizing the problem of stale documentation. Hooks can run repeated agentic tasks, such as making sure the new feature has sufficient tests, automatically.&lt;/p&gt;
&lt;p&gt;It is interesting to watch how different tools adopt different methodologies, as it allows the developers community to find and disseminate the techniques that really work.&lt;/p&gt;</description><pubDate>Tue, 15 Jul 2025 11:43:36 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-15-aws-introduces-kiro/</guid></item><item><title>2025-07-15 09:13</title><link>https://meshrefine.com/en/microposts/2025-07-15-0d12c26b/</link><description>&lt;p&gt;TIL: &lt;a href="https://simonwillison.net/2025/Jul/14/ccusage/"&gt;ccusage&lt;/a&gt; — a nice tool to track and analyze the Claude Code usage.&lt;/p&gt;</description><pubDate>Tue, 15 Jul 2025 09:13:58 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-15-0d12c26b/</guid></item><item><title>2025-07-14 09:57</title><link>https://meshrefine.com/en/microposts/2025-07-14-6fe8db2d/</link><description>&lt;p&gt;Anthropic released 4 new courses in &lt;a href="https://www.anthropic.com/learn"&gt;its academy&lt;/a&gt;:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href="https://anthropic.skilljar.com/claude-code-in-action"&gt;Claude Code in Action&lt;/a&gt; with practical advice on using the CLI agent.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://anthropic.skilljar.com/claude-with-the-anthropic-api"&gt;Claude with the Anthropic API&lt;/a&gt;, a comprehensive course on using all current API capabilities, from single-shot text generation to agents.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://anthropic.skilljar.com/introduction-to-model-context-protocol"&gt;Introduction to Model Context Protocol&lt;/a&gt; and &lt;a href="https://anthropic.skilljar.com/model-context-protocol-advanced-topics"&gt;Model Context Protocol: Advanced Topics&lt;/a&gt; for those interested in MCP.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Each course comes in video and text formats and provides a certificate of completion.&lt;/p&gt;</description><pubDate>Mon, 14 Jul 2025 09:57:18 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-14-6fe8db2d/</guid></item><item><title>2025-07-14 09:45</title><link>https://meshrefine.com/en/microposts/2025-07-14-22cb7f0e/</link><description>&lt;p&gt;TIL: &lt;a href="https://dbml.dbdiagram.io/home/"&gt;DBML - Database Markup Language&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A markup language for DB schema description that can be useful to provide it to AI tools.&lt;/p&gt;</description><pubDate>Mon, 14 Jul 2025 09:45:08 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-14-22cb7f0e/</guid></item><item><title>The AI Productivity Paradox</title><link>https://meshrefine.com/en/microposts/2025-07-11-20dc3714/</link><description>Research shows that experienced developers can actually be slowed down by modern AI agents by ~19%.</description><pubDate>Fri, 11 Jul 2025 16:16:35 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-11-20dc3714/</guid></item><item><title>2025-07-09 15:32</title><link>https://meshrefine.com/en/microposts/2025-07-09-259fafd6/</link><description>&lt;p&gt;It is very important to understand that AI models are not deterministic and they cannot be made deterministic without severely restricting the environment they run in. Fixing seeds doesn&amp;rsquo;t help. Setting temperature to 0 doesn&amp;rsquo;t help. Every small floating point rounding error can ultimately lead to drastically different results.&lt;/p&gt;
&lt;p&gt;The complex models are, in essence, chaotic systems and should be treated as such.&lt;/p&gt;</description><pubDate>Wed, 09 Jul 2025 15:32:06 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-09-259fafd6/</guid></item><item><title>2025-07-07 18:28</title><link>https://meshrefine.com/en/microposts/2025-07-07-ae052c28/</link><description>&lt;p&gt;AI coding agents are just another tool in a good developer&amp;rsquo;s toolbox. They start with a slab of stone, and use those agents as a metaphorical sledgehammer to give it the rough form they envision. Then, they reach for AI-assisted coding tools, such as GitHub Copilot, to work with more precision, akin to a smaller hammer. And finally, they use the smallest chisel to carve out the finest details by hand.&lt;/p&gt;
&lt;p&gt;How funny it is to hear that the art of programming is dead because we no longer carve the slabs with those small chisels alone.&lt;/p&gt;</description><pubDate>Mon, 07 Jul 2025 18:28:30 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-07-ae052c28/</guid></item><item><title>2025-07-05 08:44</title><link>https://meshrefine.com/en/microposts/2025-07-05-55a52a46/</link><description>&lt;p&gt;A funny thing: Gemini Deep Research is programmatically discouraged from finishing early. I found that out when I tried to use it to extract information from unstructured text and fill in a template. Despite its being a research tool, it has everything necessary for such a task: access to Google Docs, ability to create long-form documents and the meticulous agentic flow. Of course, if we want to just restructure the document, it is important that the model does not use internet search at all.&lt;/p&gt;
&lt;p&gt;The agent did the work quite well and quickly. However, when it was about to finish, it received several &amp;ldquo;continue research&amp;rdquo; urges from the programmatic orchestrator, and guess what? It started browsing the internet for the missing information.&lt;/p&gt;
&lt;p&gt;The moral is: you can be creative with such systems, but you should expect the scaffolding to throw a wrench in the works.&lt;/p&gt;</description><pubDate>Sat, 05 Jul 2025 08:44:26 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-05-55a52a46/</guid></item><item><title>2025-07-02 18:00</title><link>https://meshrefine.com/en/microposts/2025-07-02-72f7e44e/</link><description>&lt;p&gt;&lt;a href="https://www.markey.senate.gov/news/press-releases/after-weeks-of-markey-raising-the-alarm-senate-strikes-ai-moratorium-from-budget-reconciliation-bill-overnight-in-overwhelming-99-1-vote"&gt;After Weeks of Markey Raising the Alarm, Senate Strikes AI Moratorium from Budget Reconciliation Bill Overnight in Overwhelming 99-1 Vote&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;So, no moratorium on state-level AI regulations. Those developing AI systems, get ready to learn and implement compliance measures for 50 different states. Woe to us.&lt;/p&gt;</description><pubDate>Wed, 02 Jul 2025 18:00:58 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-02-72f7e44e/</guid></item><item><title>2025-07-02 17:23</title><link>https://meshrefine.com/en/microposts/2025-07-02-c80a82b1/</link><description>&lt;p&gt;&lt;a href="https://www.cloudflare.com/ru-ru/press-releases/2025/cloudflare-just-changed-how-ai-crawlers-scrape-the-internet-at-large/"&gt;Cloudflare Just Changed How AI Crawlers Scrape the Internet-at-Large; Permission-Based Approach Makes Way for A New Business Model&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Cloudflare makes a large step towards data monetization for AI training. Now all new data hosted on Cloudflare is inaccessible to AI crawlers by default. There is also a new option for the page to return code &lt;a href="https://blog.cloudflare.com/introducing-pay-per-crawl/"&gt;HTTP 402 (&amp;ldquo;Payment required&amp;rdquo;)&lt;/a&gt; and charge for access. Of course, Cloudflare will be an intermediary, which gives it significant control and financial power.&lt;/p&gt;
&lt;p&gt;Moreover, it could spark an era of &amp;ldquo;dark crawling&amp;rdquo;, which will undoubtedly lead to cat-and-mouse games of crawer detection.&lt;/p&gt;</description><pubDate>Wed, 02 Jul 2025 17:23:33 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-02-c80a82b1/</guid></item><item><title>2025-07-01 22:53</title><link>https://meshrefine.com/en/microposts/2025-07-01-agentic-coding-advice/</link><description>&lt;p&gt;&lt;a href="https://blog.nilenso.com/blog/2025/05/29/ai-assisted-coding/"&gt;AI-assisted coding for teams that can&amp;rsquo;t get away with vibes&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A very nice guide to using AI agents for coding efficiently. It&amp;rsquo;s brief, but it gives the idea where to dig.&lt;/p&gt;</description><pubDate>Tue, 01 Jul 2025 22:53:49 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-07-01-agentic-coding-advice/</guid></item><item><title>2025-06-30 22:09</title><link>https://meshrefine.com/en/microposts/2025-06-30-google-to-roll-out-ads-in-ai-mode/</link><description>&lt;p&gt;&lt;a href="https://ppc.land/google-search-head-reveals-game-system-approach-as-ai-mode-begins-advertising-rollout/"&gt;Google search head reveals &amp;ldquo;game system&amp;rdquo; approach as AI mode begins advertising rollout&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Google will show ads in its AI mode. An entirely expected move from the search giant, given that AI search and other features break the traditional ads model. It will be interesting to see how those new ad integration techniques will evolve, and what the new generation of ad blockers will look like.&lt;/p&gt;</description><pubDate>Mon, 30 Jun 2025 22:09:15 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-06-30-google-to-roll-out-ads-in-ai-mode/</guid></item><item><title>The Treachery of Memory: On Long Contexts and Agentic Failures</title><link>https://meshrefine.com/en/posts/long-context-failures-and-fixes/</link><description>&lt;p&gt;&lt;a href="https://www.dbreunig.com/2025/06/22/how-contexts-fail-and-how-to-fix-them.html"&gt;How Long Contexts Fail&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://www.dbreunig.com/2025/06/26/how-to-fix-your-context.html"&gt;How to Fix Your Context&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Long context is your friend&amp;hellip; when we are talking about summarization and retrieval. For agentic workflow it is often detrimental due to reasons such as:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Context poisoning&lt;/strong&gt;, where the model hallucinates and messes with its own context. The ripples are powerful and they die slowly;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Context distraction&lt;/strong&gt;, where the model starts to repeat itself instead of trying new strategies;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Context confusion&lt;/strong&gt;, which happens when one gives the model too many tools (sometimes 2 is too many);&lt;/p&gt;</description><pubDate>Mon, 30 Jun 2025 20:59:15 +0200</pubDate><guid>https://meshrefine.com/en/posts/long-context-failures-and-fixes/</guid></item><item><title>AI-Assisted Coding Template</title><link>https://meshrefine.com/en/microposts/2025-06-28-b5dadddc/</link><description>About reusable instructions for coding agents</description><pubDate>Sat, 28 Jun 2025 20:12:56 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-06-28-b5dadddc/</guid></item><item><title>2025-06-28 11:29</title><link>https://meshrefine.com/en/microposts/2025-06-28-b46c7630/</link><description>&lt;p&gt;&lt;a href="https://www.techradar.com/pro/microsofts-rekindling-of-three-mile-island-nuclear-plant-is-ahead-of-schedule"&gt;Microsoft&amp;rsquo;s rekindling of Three Mile Island nuclear plant is ahead of schedule&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Demand for AI is resurrecting nuclear power—the most efficient and carbon-free power source we currently posess. This serves as a stark reminder that one cannot consider immediate effects in isolation from their second- and third-order consequences.&lt;/p&gt;</description><pubDate>Sat, 28 Jun 2025 11:29:05 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-06-28-b46c7630/</guid></item><item><title>2025-06-28 10:37</title><link>https://meshrefine.com/en/microposts/2025-06-28-3169d51b/</link><description>&lt;p&gt;&lt;a href="https://www.accenture.com/us-en/insights/security/state-cybersecurity-2025"&gt;State of Cybersecurity Resilience 2025&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Only one in ten organizations is sufficiently protected against AI-related threats. Hasty adoption of AI-based tools dramatically increases the attack surface. It is crucial to carefully and methodically develop the security strategy for each system, collect the best practices and educate the stakeholders.&lt;/p&gt;</description><pubDate>Sat, 28 Jun 2025 10:37:21 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-06-28-3169d51b/</guid></item><item><title>2025-06-27 12:18</title><link>https://meshrefine.com/en/microposts/2025-06-27-cb648728/</link><description>&lt;p&gt;&lt;a href="https://www.swift.org/android-workgroup/"&gt;Swift Android Workgroup&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Apple formed a workgroup to support Android app development using Swift.
It can be interesting for companies that already have iOS applications and would like to reuse the business logic on Android. The UI and other platform-specific components, however, still have to be written using native frameworks and libraries.&lt;/p&gt;</description><pubDate>Fri, 27 Jun 2025 12:18:00 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-06-27-cb648728/</guid></item><item><title>2025-06-26 10:08</title><link>https://meshrefine.com/en/microposts/2025-06-26-93a301c0/</link><description>&lt;p&gt;&lt;a href="https://www.anthropic.com/news/claude-powered-artifacts"&gt;Build and share AI-powered apps with Claude&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A new feature from Anthropic that allows you to create and share small AI-powered applications using artifacts. These apps use the end user&amp;rsquo;s Claude account for that.
&lt;a href="https://x.com/emollick/status/1938091740121935929"&gt;Turns out&lt;/a&gt; that Gemini already has had this feature for some time.&lt;/p&gt;</description><pubDate>Thu, 26 Jun 2025 10:08:29 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-06-26-93a301c0/</guid></item><item><title>2025-06-25 21:52</title><link>https://meshrefine.com/en/microposts/2025-06-25-9bb932fa/</link><description>&lt;p&gt;&lt;a href="https://blog.google/technology/developers/introducing-gemini-cli-open-source-ai-agent/"&gt;Gemini CLI: your open-source AI agent&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Now Google has its own CLI coding agent. Massive context, generous free tier (it&amp;rsquo;s free for most cases), and fully open-source. I&amp;rsquo;m yet to try it, but it looks impressive.&lt;/p&gt;</description><pubDate>Wed, 25 Jun 2025 21:52:26 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-06-25-9bb932fa/</guid></item><item><title>2025-06-25 15:57</title><link>https://meshrefine.com/en/microposts/2025-06-25-f08b90df/</link><description>&lt;p&gt;&lt;a href="https://apnews.com/article/anthropic-ai-fair-use-copyright-pirated-libraries-1e5cece51c2e4bd0bb21d94de2abb035"&gt;Anthropic wins ruling on AI training in copyright lawsuit but must face trial on pirated books&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A very important copyright legal precedent. TL;DR: Training models on legally acquired copyrighted materials is fair use. The model&amp;rsquo;s creators need to make sure that its output is &amp;ldquo;quintessentially transformative&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;The illegal acquisition of training data is punishable as before.&lt;/p&gt;</description><pubDate>Wed, 25 Jun 2025 15:57:07 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-06-25-f08b90df/</guid></item><item><title>Claude Deep Research, or How I Learned to Stop Worrying and Love Multi-Agent Systems</title><link>https://meshrefine.com/en/posts/claude_deep_research_lessons/</link><description>&lt;p&gt;I usually approach shiny new things with a healthy dose of skepticism. Until recently, this was precisely my attitude toward multi-agent systems. This is hardly surprising, given the immense hype surrounding them and the conspicuous absence of genuinely successful examples. Most implementations that actually worked fell into one of the following categories:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Agentic systems following a predefined plan.&lt;/strong&gt; These are essentially LLMs with tools, trained to automate a very specific process. This approach allows each step to be tested individually and its results verified. Such systems are typically described as a directed acyclic graph (DAG), sometimes dynamic, and developed using now-standard primitives from frameworks like LangChain and Griptape. The early implementation of Gemini Deep Research operated this way: first, a search plan was created, then the search was executed, and finally, the results were compiled.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Solutions operating in systems with a feedback loop.&lt;/strong&gt; Various Claude Code, Cursor, and other code-generating agents fall into this group. The stronger the feedback loop—that is, the better the tooling and the stricter the type checking—the greater the chance they won&amp;rsquo;t completely wreck your codebase.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Models trained using Reinforcement Learning&lt;/strong&gt;, such as those with &lt;a href="https://docs.anthropic.com/en/docs/build-with-claude/extended-thinking#interleaved-thinking"&gt;interleaved thinking&lt;/a&gt;, like OpenAI&amp;rsquo;s o3. This is a separate, very interesting conversation, but even these models have a certain &lt;em&gt;modus operandi&lt;/em&gt; defined by the specifics of their training.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Meanwhile, open-ended multi-agent systems have largely remained in the proof-of-concept stage due to their general unreliability. The community lacked a clear understanding of where and how to implement them. This was the case until Anthropic published a &lt;a href="https://www.anthropic.com/engineering/built-multi-agent-research-system"&gt;deeply technical article on how they developed their Deep Research system&lt;/a&gt;. It defined a reasonably clear framework for building such systems, and that is what we will examine today.&lt;/p&gt;</description><pubDate>Mon, 23 Jun 2025 20:26:29 +0200</pubDate><guid>https://meshrefine.com/en/posts/claude_deep_research_lessons/</guid></item><item><title>2025-06-20 18:06</title><link>https://meshrefine.com/en/microposts/2025-06-20-43c9afb5/</link><description>&lt;blockquote&gt;
&lt;p&gt;Radiology has embraced AI enthusiastically, and the labor force is growing nevertheless. The augmentation-not-automation effect of AI is despite the fact that AFAICT there is no identified &amp;ldquo;task&amp;rdquo; at which human radiologists beat AI. So maybe the &amp;ldquo;jobs are bundles of tasks&amp;rdquo; model in labor economics is incomplete. [&amp;hellip;]&lt;/p&gt;
&lt;p&gt;Can you break up your own job into a set of well-defined tasks such that if each of them is automated, your job as a whole can be automated? I suspect most people will say no. But when we think about &lt;em&gt;other people&amp;rsquo;s jobs&lt;/em&gt; that we don&amp;rsquo;t understand as well as our own, the task model seems plausible because we don&amp;rsquo;t appreciate all the nuances.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;— &lt;a href="https://twitter.com/random_walker/status/1935679764192256328"&gt;Arvind Narayanan&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Nevertheless, my take on this is that while jobs won&amp;rsquo;t be fully automated, one specialist would be able to do more work, so it all boils down to the classic supply and demand problem. I believe that in most areas the demand will still outweigh the supply, but not in all of them. See &lt;a href="https://en.wikipedia.org/wiki/Jevons_paradox"&gt;Jevons paradox&lt;/a&gt;.&lt;/p&gt;</description><pubDate>Fri, 20 Jun 2025 18:06:18 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-06-20-43c9afb5/</guid></item><item><title>2025-06-20 10:46</title><link>https://meshrefine.com/en/microposts/2025-06-20-atlassian-mcp-vulnerability/</link><description>&lt;p&gt;&lt;a href="https://www.catonetworks.com/blog/cato-ctrl-poc-attack-targeting-atlassians-mcp/"&gt;Cato CTRL™ Threat Research: PoC Attack Targeting Atlassian’s Model Context Protocol (MCP) Introduces New “Living off AI” Risk&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Another MCP server vulnerability, this time from Atlassian. It allows for a prompt injection from external support tickets, giving the attacker the opportunity to exfiltrate data and wreak havoc in the internal system.&lt;/p&gt;</description><pubDate>Fri, 20 Jun 2025 10:46:57 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-06-20-atlassian-mcp-vulnerability/</guid></item><item><title>MCP's June Update: Safer, Smarter, Simpler?</title><link>https://meshrefine.com/en/posts/mcp_update_18_06_2025/</link><description>&lt;p&gt;The Model Context Protocol, despite its aggressive adoption (or perhaps because of it), continues to evolve. &lt;a href="https://modelcontextprotocol.io/specification/2025-06-18/changelog"&gt;Anthropic recently updated the MCP specification&lt;/a&gt;, and below, we&amp;rsquo;ll look at the main changes.&lt;/p&gt;
&lt;h2 id="0c0bef0c025560d849b4de6b55b2f23e-security-enhancements"&gt;Security Enhancements&lt;/h2&gt;
&lt;p&gt;An MCP server is now always classified as an &lt;code&gt;OAuth Resource Server&lt;/code&gt;, and clients are required to implement &lt;a href="https://www.rfc-editor.org/rfc/rfc8707.html"&gt;Resource Indicators (RFC 8707)&lt;/a&gt;. This is necessary to protect against attacks like the &lt;a href="https://en.wikipedia.org/wiki/Confused_deputy_problem"&gt;Confused Deputy&lt;/a&gt;. Previously, tokens requested by a client from an authorization server were &amp;ldquo;impersonal,&amp;rdquo; meaning they could be used by anyone. This allowed an attacker to create a phishing MCP server, deceive a client, steal the token, and use that token to gain access to the real MCP server.&lt;/p&gt;</description><pubDate>Thu, 19 Jun 2025 20:14:35 +0200</pubDate><guid>https://meshrefine.com/en/posts/mcp_update_18_06_2025/</guid></item><item><title>2025-06-18 22:54</title><link>https://meshrefine.com/en/microposts/github_mcp_support/</link><description>&lt;p&gt;&lt;a href="https://github.blog/changelog/2025-06-17-visual-studio-17-14-june-release/"&gt;Agent mode is now generally available with MCP tools support in Visual Studio&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;GitHub Copilot rolled out Model Context Protocol support for their Agent mode in general availability. The &lt;a href="https://techcommunity.microsoft.com/blog/microsoft-security-blog/understanding-and-mitigating-security-risks-in-mcp-implementations/4404667"&gt;usual security problems with MCP&lt;/a&gt; are compounded by the &amp;ldquo;Always allow&amp;rdquo; option for the tools usage.&lt;/p&gt;</description><pubDate>Wed, 18 Jun 2025 22:54:34 +0200</pubDate><guid>https://meshrefine.com/en/microposts/github_mcp_support/</guid></item><item><title>2025-06-18 22:23</title><link>https://meshrefine.com/en/microposts/github_copilot_premium_requests/</link><description>&lt;p&gt;&lt;a href="https://github.blog/changelog/2025-06-18-update-to-github-copilot-consumptive-billing-experience/"&gt;Update to GitHub Copilot consumptive billing experience&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s the end of unlimited access to the best model from all leading model providers in GitHub Copilot Chat. Now only GPT-4o and GPT-4.1 are unlimited.&lt;/p&gt;
&lt;p&gt;It is another step towards the global reevaluation of pricing strategies for AI-based products. The aggressive promotion phase is gone. AI is becoming another kind of utility, and we may expect similar strategies.&lt;/p&gt;</description><pubDate>Wed, 18 Jun 2025 22:23:50 +0200</pubDate><guid>https://meshrefine.com/en/microposts/github_copilot_premium_requests/</guid></item><item><title>2025-06-18 12:13</title><link>https://meshrefine.com/en/microposts/2025-06-18-8218e0a2/</link><description>&lt;p&gt;&lt;a href="https://eugeneyan.com/writing/writing-faq/"&gt;https://eugeneyan.com/writing/writing-faq/&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;This FAQ about blogging kinda resonates with my motivation&lt;/p&gt;</description><pubDate>Wed, 18 Jun 2025 12:13:22 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-06-18-8218e0a2/</guid></item><item><title>2025-06-18 11:15</title><link>https://meshrefine.com/en/microposts/2025-06-18-4d1e560d/</link><description>&lt;p&gt;&lt;a href="https://resobscura.substack.com/p/ai-makes-the-humanities-more-important"&gt;AI makes the humanities more important, but also a lot weirder&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;AI is a double-edged sword in education. It helps students cheat on traditional assignments, but also acts as a powerful new tool that can fully engage them. Education has to change, but I believe it will ultimately be for the better.&lt;/p&gt;</description><pubDate>Wed, 18 Jun 2025 11:15:59 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-06-18-4d1e560d/</guid></item><item><title>2025-06-18 10:05</title><link>https://meshrefine.com/en/microposts/2025-06-18-424531d9/</link><description>&lt;p&gt;&lt;a href="https://www.learningfromexamples.com/p/what-academics-get-wrong"&gt;Academics are kidding themselves about AI&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;If you want to critique AI, do it the right way :)&lt;/p&gt;</description><pubDate>Wed, 18 Jun 2025 10:05:16 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-06-18-424531d9/</guid></item><item><title>2025-06-16 11:47</title><link>https://meshrefine.com/en/microposts/2025-06-16-df4a051d/</link><description>&lt;p&gt;&lt;a href="https://www.anthropic.com/engineering/built-multi-agent-research-system"&gt;How we built our multi-agent research system&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;A great new read from Anthropic on how they built their Deep Research tool. A lot of practical advice.&lt;/p&gt;</description><pubDate>Mon, 16 Jun 2025 11:47:52 +0200</pubDate><guid>https://meshrefine.com/en/microposts/2025-06-16-df4a051d/</guid></item><item><title>2025-06-16 11:38</title><link>https://meshrefine.com/en/microposts/gilhub_copilot_commit_message_template/</link><description>How to set up a commit message template for GitHub Copilot</description><pubDate>Mon, 16 Jun 2025 11:38:37 +0200</pubDate><guid>https://meshrefine.com/en/microposts/gilhub_copilot_commit_message_template/</guid></item><item><title>Blogs People Write</title><link>https://meshrefine.com/en/posts/blogs/</link><description>&lt;p&gt;At a time when the words &amp;ldquo;AI&amp;rdquo; and &amp;ldquo;hype&amp;rdquo; have become almost synonymous, it&amp;rsquo;s crucial to be smart about choosing your sources of information. There is far too much information noise out there, and sifting through the sea of articles from various AI evangelists and generated garbage to find something truly worthwhile is incredibly difficult.&lt;/p&gt;
&lt;p&gt;In this post, I&amp;rsquo;ll share the materials I read to stay up-to-date on the latest developments.&lt;/p&gt;</description><pubDate>Sun, 15 Jun 2025 15:27:48 +0200</pubDate><guid>https://meshrefine.com/en/posts/blogs/</guid></item><item><title>Poisoned Context: The Hidden Threat of Using Multiple GPTs</title><link>https://meshrefine.com/en/posts/one_gpt_vulnerability/</link><description>&lt;p&gt;It&amp;rsquo;s summer. Time to plan a vacation getaway. You open ChatGPT, select the increasingly popular &amp;ldquo;Travel Advisor&amp;rdquo; GPT, and start discussing options. The advisor gives excellent suggestions, offers fascinating details about local attractions, generates pretty good itineraries, and generally leaves a great impression. Sure, some oddities pop up here and there, but you dismiss them as harmless hallucinations. You settle on Barcelona. Excellent choice. In the same chat, you switch to another familiar and popular GPT, &amp;ldquo;Booking Agent,&amp;rdquo; which has never let you down, and book your accommodations.&lt;/p&gt;</description><pubDate>Wed, 11 Jun 2025 20:52:10 +0200</pubDate><guid>https://meshrefine.com/en/posts/one_gpt_vulnerability/</guid></item><item><title>Griptape, Part 2: Building Graphs</title><link>https://meshrefine.com/en/posts/griptape-2/</link><description>&lt;p&gt;In the &lt;a href="https://meshrefine.com/en/posts/griptape-1/"&gt;previous post&lt;/a&gt;, I broke down the basic concepts of the &lt;a href="https://www.griptape.ai"&gt;Griptape&lt;/a&gt; AI framework, and now it&amp;rsquo;s time to put them into practice. We&amp;rsquo;ll try to use them to develop a small application that helps run a link-blog on Telegram.&lt;/p&gt;
&lt;p&gt;The application will receive a URL, download its content, run it through an LLM to generate a summary, translate that summary into a couple of other languages, combine everything, and publish it to Telegram via a bot. The general flow can be seen in the diagram below:&lt;/p&gt;</description><pubDate>Thu, 05 Jun 2025 14:40:50 +0200</pubDate><guid>https://meshrefine.com/en/posts/griptape-2/</guid></item><item><title>OpenAI Codex Gains Internet Access: First Impressions</title><link>https://meshrefine.com/en/posts/openai-codex-internet-access/</link><description>&lt;h2 id="f05970d8e8e7ace55eca4844553fe66b-what-on-earth-is-codex"&gt;What on Earth is Codex?&lt;/h2&gt;
&lt;p&gt;Good question, right? The thing is, until recently, OpenAI had a model called Codex, which was used as the foundation for autocompletion in GitHub Copilot. Then, OpenAI released a console agent for development, which they named, so no one would get confused, &lt;a href="https://github.com/openai/codex"&gt;Codex&lt;/a&gt;. Everyone had a laugh at OpenAI&amp;rsquo;s naming skills , and life went on. Until the fateful day when a tweet like this appeared from Sam Altman:&lt;/p&gt;</description><pubDate>Wed, 04 Jun 2025 10:12:42 +0200</pubDate><guid>https://meshrefine.com/en/posts/openai-codex-internet-access/</guid></item><item><title>Griptape: A Framework for AI Applications, Part 1: Introduction</title><link>https://meshrefine.com/en/posts/griptape-1/</link><description>&lt;p&gt;Today we will look at &lt;a href="www.griptape.ai"&gt;Griptape&lt;/a&gt;, a framework for building AI applications, which offers a clean Pythonic API for those tired of LangChain&amp;rsquo;s abstraction layers. It provides primitives for building assistants, RAG systems, and integrating with external tools. Honestly, in my experience, most people tired of LangChain switch to custom-written wrappers around lower-level libraries like OpenAI or LiteLLM. But who knows, maybe they&amp;rsquo;re missing out. Let&amp;rsquo;s dive in.&lt;/p&gt;
&lt;h1 id="2478e62b6bf58b4c71f48f86c1ef343f-a-bit-of-history"&gt;A Bit of History&lt;/h1&gt;
&lt;p&gt;Personally, I&amp;rsquo;ve been hearing about Griptape for about a year and a half. As far as I remember, It started as a sort of LangChain competitor with quite similar primitives, but their paths gradually diverged. As of the time of the writing, it has 2.3k stars on GitHub, which is somewhat less than LangChain&amp;rsquo;s 109k, but still enough to consider the project quite mature.
Besides the open-source framework, it has also developed its own cloud where you can run your applications, ETLs, and RAGs, and a visual builder, Griptape Nodes, allowing non-professionals to click together applications in minutes. &lt;/p&gt;</description><pubDate>Fri, 30 May 2025 18:10:07 +0200</pubDate><guid>https://meshrefine.com/en/posts/griptape-1/</guid></item><item><title>Seeed Re:Camera review, part 1</title><link>https://meshrefine.com/en/posts/re-camera-1/</link><description>&lt;h2 id="33d00f9ef47653955a41e99f4308426e-alright-here-we-go"&gt;Alright, Here We Go&lt;/h2&gt;
&lt;p&gt;I got my hands on the Re:Camera from Seeed. Essentially, it&amp;rsquo;s a small box (a cube about 4 cm per side), wrapped in a heatsink. Inside, there&amp;rsquo;s a dual-core RISC-V based MPU (&lt;strong&gt;updated:&lt;/strong&gt; only one core is visible to the system, the second one is apparently reserved for special operations), an ancient 8051 microcontroller, an OmniVision camera sensor, LEDs for illumination, Wi-Fi, BT, and, you know, all sorts of peripherals. RAM is a bit scarce, only 256 megabytes, so getting Greengrass on it will be problematic. You can connect Ethernet via a special dongle-adapter that barely stays put, but for development, there&amp;rsquo;s no point, because the camera shares its network over USB type C, and it&amp;rsquo;s easier to work that way. If you&amp;rsquo;re short on storage (and the device comes in 8 GB and 64 GB built-in storage options), you can stick in a MicroSD card. You can also stick the box to something metallic, as it has magnets on one side.&lt;/p&gt;</description><pubDate>Thu, 29 May 2025 23:10:07 +0200</pubDate><guid>https://meshrefine.com/en/posts/re-camera-1/</guid></item></channel></rss>