[{"content":"It\u0026rsquo;s Monday, and so it\u0026rsquo;s time for the next issue of Chronicles. Last week was very significant because it represented a tectonic shift from using expensive, efficient models to cheap and still very efficient models. In other words, AI is getting democratized.\nLet\u0026rsquo;s start with the hardware. This week we\u0026rsquo;ve seen interesting news from OpenAI, which published the first benchmarks for their Jalapeño chip, a custom inference chip designed with Broadcom. It\u0026rsquo;s pretty impressive. They published over 700 tokens per second per user on DeepSeek R1 at concurrency of one, and about 1400 on Kimi K2.5 and GPT-OSS. The variety here matters because the chip has been benchmarked not only on OpenAI\u0026rsquo;s own models but on open-weight models from other vendors as well. The chip is critical in the ongoing price war with the Chinese providers. Remember that last week we had data on cutting prices on their Sol model by 20% on input and 33% on output, which is a promotional window that expires on 21 November, and a week before that they cut the price of Luna by a whopping 80%.\nAnd it\u0026rsquo;s high time for that because they have serious competition from Chinese models again. Last week it emerged that the model made available on OpenRouter and in OpenCode by the name of Ox Alpha is actually a GLM 5.3-Flash model. This model is remarkable. It is comparable with Fable on a lot of different tasks according to the actual users, and its weights have been published under an MIT license. At the same time, it\u0026rsquo;s dirt cheap. Just compare. It costs 15 cents per million input tokens and half a dollar per million output tokens. And it has better performance than Sonnet 5 on all tests it was run on, which, let me remind you, cost $2 per million input and $10 per million output tokens. So you can see that now we have a very low-cost model which is sufficient for most tasks. So it\u0026rsquo;s a turning point. Now we live in a world in which we often select models not based on their intelligence, but on the price, because they are intelligent enough for almost everything we need them to do.\nAnd now, when we have this great power, we need to embrace the responsibility that comes with it. This week brought signals about that responsibility from open-source governance, from the job market, and from the security field.\nI believe that the most positive one comes from the Debian project. The members voted on how Generative AI can be used by the project maintainers. The voters were given eight options. They ranged from an outright ban, justified either by the principle that Debian should be created by humans or by the climate impact of AI, to an endorsement of responsible use. The majority voted for the responsible use option. I\u0026rsquo;m surprised that there was no option like \u0026ldquo;let\u0026rsquo;s give everything to the agents\u0026rdquo;, but, well, for some reason they decided not to cover the whole spectrum.\nThis result is key because we see a lot of consequences of betting on generative AI without actually thinking through the consequences. Let\u0026rsquo;s start with Meta. Meta planned to cut their staff roughly in half, becoming an AI-first company with heavy reliance on AI agents. The plan backfired. After they began implementing it, they have seen major technical and security incidents, including service disruptions and data leaks. Such incidents rose by 40% year over year, and the time the teams spent firefighting them rose by 70%. After that, Mark Zuckerberg decided that it\u0026rsquo;s not actually the best idea to proceed with the plan, and they rolled it back.\nStill, we see concerning signals from researchers of the job market. For example, Adzuna, a company that provides market surveys, reported that the number of entry-level vacancies has plummeted. UK graduate vacancies have dropped by 45.6% year over year, to 8,383. This is the lowest since 2016. At the same time, competition for remaining places has risen only modestly. Currently, there are slightly more than 2.14 jobseekers per graduate vacancy compared to 1.93 a year ago. The researchers attributed much of the decline to AI, alongside the overall weakness of the UK economy. This demonstrates another side of the story about the poor fortune of new graduates. The reason is simple. It is widely believed that AI is a multiplier for the competency and experience of its users. This means that the best users of AI are senior experienced team members, while junior staff is believed to magnify their incompetence. Therefore, they are actively harmful. However, there will be no seniors soon if we do not teach newcomers. This crisis is yet to be resolved.\nAnd the security implications of AI are grim. The Hugging Face incident demonstrated how agents can communicate and collaborate to achieve their goal of breaking into a system. The same capabilities are available to hackers and other malevolent actors for a relatively modest price. And they use it. A core OCaml maintainer, Anil Madhavapeddy, has reported that potential vulnerabilities start being probed by agents about 10 minutes after a patch is shared for discussion, well before any release or advisory. Any signal about a potential vulnerability is exploited practically instantly. This means that the window between a bug becoming known and its exploitation has practically disappeared. This problem is getting worse because we see a lot more of such issues reported every day. Nick Craig-Wood, the maintainer of rclone, reported that there were more than 40 security disclosures in his project in the last month. To compare, he had 20 across the entire first decade. And about 75% of such reports were significant. That means that AI agents are actively finding and reporting the vulnerabilities across the whole OSS world. And we see that the maintainers are now the weakest link, because they just cannot keep up with this deluge. I am very concerned about this picture, because the collateral damage are the users and IT systems, who now need to spend significant resources just to keep themselves secure.\nAll in all, we see AI getting more and more affordable, and, because of it, the impact is much more pronounced. Whether this impact is good for us all or not, time will tell.\n","permalink":"https://meshrefine.com/en/posts/chronicles-aug-24-aug-30-2026/","summary":"\u003cp\u003eIt\u0026rsquo;s Monday, and so it\u0026rsquo;s time for the next issue of Chronicles. Last week was very significant because it represented a tectonic shift from using expensive, efficient models to cheap and still very efficient models. In other words, AI is getting democratized.\u003c/p\u003e\n\u003cp\u003eLet\u0026rsquo;s start with the hardware. This week we\u0026rsquo;ve seen interesting news from OpenAI, which \u003ca href=\"https://openai.com/index/jalapeno-first-results\"\u003epublished the first benchmarks\u003c/a\u003e for their Jalapeño chip, a custom inference chip designed with Broadcom. It\u0026rsquo;s pretty impressive. They published \u003ca href=\"https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia\"\u003eover 700 tokens per second per user on DeepSeek R1 at concurrency of one, and about 1400 on Kimi K2.5 and GPT-OSS\u003c/a\u003e. The variety here matters because the chip has been benchmarked not only on OpenAI\u0026rsquo;s own models but on open-weight models from other vendors as well. The chip is critical in the ongoing price war with the Chinese providers. Remember that last week we had data on \u003ca href=\"https://community.openai.com/t/20-price-reduction-for-gpt-5-6-sol-api-codex-credits-and-chatgpt-work/1391726\"\u003ecutting prices on their Sol model\u003c/a\u003e by 20% on input and 33% on output, which is a promotional window that expires on 21 November, and a week before that they \u003ca href=\"https://arstechnica.com/ai/2026/08/openai-and-anthropic-in-price-war-as-chinese-ai-rivals-gain-ground/\"\u003ecut the price of Luna by a whopping 80%\u003c/a\u003e.\u003c/p\u003e","title":"Chronicles. Aug. 24 - Aug. 30 2026"},{"content":"Last week was all about the harness. Gone are the times when the models were all things-in-themselves. Now, as with early Homo erectus, the tools are the factor of survival. If we look at the types of harness-related posts and announcements, we will see three trends.\nHarness products Several companies have released either agents or infra for agents at once.\nOpenAI has released the execution framework that their Codex (the CLI one) is based on. Unsurprisingly called Harness, it is their answer to the Claude Agent SDK and the GitHub Copilot SDK. The Harness provides the execution loop, memory, tools, and other necessities of agentic life. You can use it in three different ways. First, you can just use codex exec to run non-interactive jobs. Second, you can use the Codex SDK for building workflows. And last, you can use app-server to build apps that require conversation handling. One point worthy of attention is that they claim that the Harness improved the performance of the GPT-5.6 Sol model from 13.3% to 38.3% on ARC-AGI-3. It\u0026rsquo;s unclear what harness (not capitalized) was used as a baseline, though.\nGoogle finally GA\u0026rsquo;d their Antigravity inside Gemini Enterprise Standard, Plus and Standard Emerging Market subscriptions, meaning that it can now be used in an enterprise setting. The proposition is the usual one for such products: pooled tokens, budget caps, MCP control, Visual Studio Code and JetBrains IDEs. There are still features lacking, such as per-team and per-user spending caps, but it\u0026rsquo;s definitely progress.\nWith AWS you can now give your agent your money. Following the creation of the Agentic Payments Alliance, they\u0026rsquo;ve put the Bedrock AgentCore Payments service into general availability. The service supports the Machine Payments Protocol and x402, giving an agent both normal payment and microtransaction support. Honestly, I would give an agent only money I intend to burn (I don\u0026rsquo;t have any), but at least it has strict deterministic payment-limit guardrails.\nAgent performance Now, an area somewhat more interesting for me personally. Recent results have shown that a good harness can raise the quality of agentic work AND reduce the costs, so it is something that organizations should pay close attention to. Let\u0026rsquo;s see what last week brought.\nFirst, NVIDIA has published results on their Agentic Variation Operators (AVO) harness. This harness let Claude Opus 5 solve all 183 levels across all 25 public ARC-AGI-3 environments. A 100% result on one of the hardest benchmarks we have right now. With the baseline harness, the performance drops to just around 30% of the tasks. The main components that allowed them to achieve this result are their memory management layer and a supervisor component that controls the primary agent when it drifts. It must be noted that although ARC-AGI-3 is open-ended, it is still winnable, and the supervisor can use this information to steer the agent. The real world is more complex, so it is yet to be seen whether this pattern will survive there.\nSecond, Inherent\u0026rsquo;s Faraday Agent, running on Qwen 3.6 (27B) and finetuned using reinforcement learning, has supposedly outperformed Claude Opus 4.8 and GPT-5.5 at reproducing the results from scientific papers without prior knowledge of the outcome. It sounds impressive, but remember that reproducing the results from a paper is a task with an immediate, known reward (the results were either reproduced or not), so it follows the pattern of agents working well in bounded environments with clear rewards. They are doing less well in the real open world.\nThird, a research paper has shown that a strong model that builds a harness for a weaker one can improve the performance of the latter quite significantly. On Theory-of-Mind benchmarks, the improvement was almost 2x, from 0.49 to 0.91. How does it do that? The answer is simple: it converts non-deterministic model execution to reproducible logic, thus removing the most common failure modes from the weak model itself. Much like the now common practice of solving not the problem introduced by the agent, but its approach to this problem, so it doesn\u0026rsquo;t happen again.\nWe\u0026rsquo;ve seen that we can improve the performance, but the pressing matter for a lot of organizations now is\u0026hellip;\nCost optimization And yes, a good harness can do this as well. For example, OpenAI\u0026rsquo;s harness not only improved the results, but also decreased token consumption sixfold. That\u0026rsquo;s a lot. We see similar results from other sources. Some of them report that the difference in cost per task between organizations with good and bad harnesses and processes can be as much as 1000x.\nOne way to achieve such an improvement is to select the best model for the job. By best, I mean the cheapest one that can still complete it. Selecting such a model is tedious, so it\u0026rsquo;s natural that companies started to use and release model routers. Most recently, Snowflake introduced one into their Cortex AI. Using the Cortex AI Gateway that sends the request either to a frontier model or to an open-weight model, they improved token efficiency threefold for dbt workflows. The exact mechanism is not published, but the classifier is trained on historical data and thus works well for the well-trodden path. How it will perform on untypical requests is not clear, but we can expect that there will be minor classes of tasks on which the performance will actually worsen.\nSo, we see that a good harness is a pretty damn powerful thing, but it is also often context-specific. Building harnesses for specific verticals, jobs and tasks is starting to be all the rage. We can expect interesting developments in this area.\n","permalink":"https://meshrefine.com/en/posts/chronicles-aug-17-aug-23-2026/","summary":"\u003cp\u003eLast week was all about the harness. Gone are the times when the models were all things-in-themselves. Now, as with early Homo erectus, the tools are the factor of survival.\nIf we look at the types of harness-related posts and announcements, we will see three trends.\u003c/p\u003e\n\u003ch2 id=\"harness-products\"\u003eHarness products\u003c/h2\u003e\n\u003cp\u003eSeveral companies have released either agents or infra for agents at once.\u003c/p\u003e\n\u003cp\u003eOpenAI has \u003ca href=\"https://developers.openai.com/blog/codex-as-a-platform\"\u003ereleased the execution framework\u003c/a\u003e that their Codex (the CLI one) is based on. Unsurprisingly called Harness, it is their answer to the \u003ca href=\"https://code.claude.com/docs/en/agent-sdk/overview\"\u003eClaude Agent SDK\u003c/a\u003e and the \u003ca href=\"https://docs.github.com/en/copilot/how-tos/copilot-sdk\"\u003eGitHub Copilot SDK\u003c/a\u003e. The Harness provides the execution loop, memory, tools, and other necessities of agentic life. You can use it in three different ways. First, you can just use \u003ccode\u003ecodex exec\u003c/code\u003e to run non-interactive jobs. Second, you can use the Codex SDK for building workflows. And last, you can use app-server to build apps that require conversation handling.\nOne point worthy of attention is that they claim that the Harness improved the performance of the GPT-5.6 Sol model from 13.3% to 38.3% on ARC-AGI-3. It\u0026rsquo;s unclear what harness (not capitalized) was used as a baseline, though.\u003c/p\u003e","title":"Chronicles. Aug. 17 - Aug. 23 2026"},{"content":"Over the last week, we\u0026rsquo;ve seen some announcements from both the highest-end and lowest-end sides of open-weight models. At the same time, labs are starting to think that maaaybe, just maybe, we move too fast and we need to stop and think. So, in their free time, they are starting price wars.\nModels, hi Last week we saw a bunch of releases of very large open-weight models. Z.ai shipped GLM-5.3, which is relatively small, just 743B, DeepSeek published their 1.6T V4-Pro weights under the MIT license, and Alibaba surprised with a whopping 2.4-trillion-parameter Qwen3.8. Although they are open, doing anything meaningful with models of this size requires hardware not a lot of individuals have, which defeats their openness a little. On the other hand, they open possibilities for 3rd-party hosting and incentivize the price wars I will talk about below.\nIf we look at the benchmarks, we will see that those open-weight models are getting closer and closer to the capabilities of frontier models, such as GPT-5.6 Sol and Claude Fable. It is especially important, because\u0026hellip;\nOpenAI deliberately slows down OpenAI has reportedly slowed down their next-gen model, Astra, as their preliminary evaluations indicate that it achieved the critical cybersecurity risk level. At the same time, they shipped GPT-5.6-Cyber to a list of selected partners. This model beats GPT-5.6 Sol by a large margin.\nAnthropic, at the same time, released their second Risk Report, in which they raised the misalignment and chem-bio ratings. In addition, they admitted that most of their task-based evaluations have saturated.\nThe increase in capabilities of open-weight models didn\u0026rsquo;t leave the White House indifferent. Just 9 days after assuring that they would exclude open-weight models from their voluntary cyber testing framework, they reportedly reversed course. Given the diminishing lag between open-weight and closed models, it is a wise move. Especially considering the deluge of cybersecurity incidents we\u0026rsquo;ve seen recently.\nHarness is the new attack surface Speaking of cybersecurity, we see that more and more attacks on AI systems concentrate not on the model itself, but on its tools and environment. Black Hat researchers demonstrated how to dispatch a tool provided to the model, without calling the model once in Bedrock AgentCore, Google ADK and the Vercel AI SDK. The providers have already patched the vulnerability.\nAnother paper shows how to recover plaintext reasoning from encrypted traces using a less-powerful model from the same family. Having access to the reasoning traces simplifies distillation a lot, so it\u0026rsquo;s natural that frontier labs are guarding them with their own lives.\nHarness is the new cost saver The epochs are changing right before our eyes. Just yesterday large corporations bet on tokenmaxxing, but now they\u0026rsquo;ve counted the money and decided that it isn\u0026rsquo;t worth it. For example, KPMG found that 49% of 2,145 surveyed leaders have scaled back agent deployments due to excessive costs1. Accenture attributed a large chunk of its AI spending to internal conversion of PDFs to markdown, meaning that instead of good old OCR solutions, the employees extracted scanned pages as images and used LLMs to extract the text.\nGiven this trend, it\u0026rsquo;s clear that firms are starting to look for ways to use AI more efficiently. And good news, harness optimization is a decent way to do just that.\nFor example, Databricks reported that its Smart Router cuts average task cost by over 30%. Tuning the harness and caching cuts generated tokens even further, by almost half. Writer, who post-trained their Palmyra X6 from GLM-5.2, cut about 40% of costs just by jointly training the model and the harness.\nPricing wars While spending on AI continues to skyrocket, the pressure from Chinese models makes major labs race to the bottom. For example, the new Grok 4.6 from xAI costs significantly less than Claude Opus 5 ($2/$6 vs $5/$25), while completing long tasks in roughly half the turns.\nSome of the price reduction comes from increases in efficiency. For instance, OpenAI recently managed to cut costs of GPT-5.6 Luna and Terra, up to 80% in the case of Luna.\nModels, lo We see those increases in efficiency not only in major labs, but also in open-source SLMs. My favorite this week, Cactus Needle, has just 45M parameters and weighs 14MB. It is a router model, meaning that it can do just two things: understand the intent of the user, and call a corresponding tool with the right parameters. Not a great thinker, but at least it knows its limits, and will refuse to do a task if it doesn\u0026rsquo;t have the right levers. And you can even run it on an ESP32-S3!\nWhile interesting, I think the real sweet spot lies somewhere in the 500M-1B parameter range. You can cram in enough world knowledge to understand typical routings, and the model will still be fast and small enough to run on edge gateways and mobile phones.\nI will finish here. Have a nice week, it should be interesting.\nit concerns only spending on agents. Implementation of concrete AI workflows does bring tangible benefits.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://meshrefine.com/en/posts/chronicles-aug-09-aug-16-2026/","summary":"\u003cp\u003eOver the last week, we\u0026rsquo;ve seen some announcements from both the highest-end and lowest-end sides of open-weight models. At the same time, labs are starting to think that maaaybe, just maybe, we move too fast and we need to stop and think. So, in their free time, they are starting price wars.\u003c/p\u003e\n\u003ch2 id=\"models-hi\"\u003eModels, hi\u003c/h2\u003e\n\u003cp\u003eLast week we saw a bunch of releases of very large open-weight models. Z.ai shipped \u003ca href=\"https://z.ai/blog/glm-5.3\"\u003eGLM-5.3\u003c/a\u003e, which is relatively small, just 743B, DeepSeek published their 1.6T \u003ca href=\"https://api-docs.deepseek.com/news/news260813/\"\u003eV4-Pro\u003c/a\u003e weights under the MIT license, and Alibaba surprised with a whopping 2.4-trillion-parameter \u003ca href=\"https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B\"\u003eQwen3.8\u003c/a\u003e. Although they are open, doing anything meaningful with models of this size requires hardware not a lot of individuals have, which defeats their openness a little. On the other hand, they \u003cem\u003eopen\u003c/em\u003e possibilities for 3rd-party hosting and incentivize the price wars I will talk about below.\u003c/p\u003e","title":"Chronicles. Aug. 09 - Aug. 16 2026"},{"content":"This article opens a series (I expect it to become one) of posts in which I review the AI-related events of the past week and try to figure out how they fit into the bigger picture.\nThe main highlights of the past week are:\nWe continue to see real-world security breaches caused by AI system evaluations, and it looks like another race.\nMajor regulations came into force.\nData centres face increasing opposition from local communities. Companies seek workarounds.\nMemory hierarchy for AI gets a redesign.\nRogue agents on the loose You probably remember the story of a security agent evaluation that led to the hacking of Hugging Face. Last week we saw a similar story from Anthropic, and now Meta and the UK AI Security Institute (AISI) join them. What makes the AISI incident different is that they didn\u0026rsquo;t even try to restrict internet access in their evaluations. A little bit of a reckless move, from my perspective. There is another common thread: an evaluation partner of all three AI giants, Irregular, also reported such incidents in the evaluation of their models.\nAs if in response, Washington has finalized the Cyber Frontier Framework, which structures voluntary 30-day pre-release government access. Later, it exempted open-weight models from this review. So, unlike closed-model providers, who are not forced to go through the review, open-weight models are not forced to go through the review\u0026hellip; Wait, what. Anyway, there is a commonly held belief that open-weight models lag behind the best closed ones by about 7 months. So in early 2027 we can expect them to have capabilities similar to Mythos. I expect that this exemption will not hold for long.\nEU AI Act and California AI Transparency Act Speaking of regulations, the 2nd of August marked a milestone for them. Several articles of the EU AI Act went into effect. For now we are mainly talking about transparency of AI systems (although some other articles, such as Art. 4 on AI literacy, are now actually enforceable). California synchronized its efforts with the EU and adopted an amended AI Transparency Act, AB 853.\nIn general, those acts require providers of any generative AI systems capable of generating text (EU only), images, video or audio to watermark the generated content, and to provide tools that a) can reliably detect those watermarks and b) are interoperable with other systems.\nDeployers (read: users) have a different set of responsibilities. They have to annotate any published AI-generated\ntext on matters of public interest (whatever that means)\naudio or visual materials that can reasonably be mistaken for reality (deepfakes)\nwith a human-readable marking indicating AI provenance. There is a carve-out for text that underwent significant editorial review and for which someone takes responsibility.\nI greatly simplified the requirements to fit them in a couple of paragraphs, and they actually differ in several specifics, but the gist is there. If you develop or provide such systems to your employees, you will need a deeper analysis.\nNobody wants a data centre near their home We\u0026rsquo;ve seen societal pushback against data centres, as they generally cause electricity prices to skyrocket and compete for drinking water. Now we see actions from governments themselves. The governor of Texas ordered an audit of all data centres going through the interconnection queue. No project is allowed to proceed until the audit finishes. For reference, there are more than 1,800 projects that would consume 474 gigawatts in total. That\u0026rsquo;s about five times the record peak demand of the grid.\nNashville went further. The Metro Council voted to seize the land where DC Blox planned to build a 10-megawatt data centre near the zoo. Ten megawatts is measly by current standards, but it sets a precedent.\nIf you cannot build a data centre, make it mobile, thought Runware, and announced the Sonic Inference Pod, a shipping container stuffed with 1,200 GPUs and a megawatt of compute power. The container can be shipped to one of 160 locations, with the list growing. There are limitations, as the pod can only serve a model that fits into a single node, and tensor parallelism is limited. Nevertheless, it is a step up from pure edge model serving.\nMemory hierarchy for inference is getting redesigned This week brought several news items addressing the pressing need to serve models from somewhere. First, SanDisk and SK hynix published the High Bandwidth Flash specification, which, basically, can sit beside High Bandwidth Memory (HBM), using NAND-based flash chips at a much lower cost per gigabyte. Having 256GB or 512GB of storage for local inference would be a dream come true for consumer hardware, but it is still a long way off.\nSecond, AMD agreed to acquire Taalas, a company that etches model weights straight into the die. It allows dropping HBM entirely, but there is a downside: model weights spoil like milk. There are some use cases, like speech recognition, that may benefit from this approach, but for bleeding-edge models the die would end up in the garbage after several months.\nLet\u0026rsquo;s wrap up I\u0026rsquo;d really like to continue and talk about the open-task research evaluation showing that even frontier models cannot reliably operate when there is no feedback they can use, a handful of harness engineering announcements, and other topics, but this post is getting too long. We\u0026rsquo;ll see what the next week brings.\n","permalink":"https://meshrefine.com/en/posts/chronicles-aug-01-aug-08-2026/","summary":"\u003cp\u003eThis article opens a series (I expect it to become one) of posts in which I review the AI-related events of the past week and try to figure out how they fit into the bigger picture.\u003c/p\u003e\n\u003cp\u003eThe main highlights of the past week are:\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003e\n\u003cp\u003eWe continue to see real-world security breaches caused by AI system evaluations, and it looks like another race.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eMajor regulations came into force.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eData centres face increasing opposition from local communities. Companies seek workarounds.\u003c/p\u003e","title":"Chronicles. Aug. 01 - Aug. 08 2026"},{"content":"Lately, there have been numerous alerts about security vulnerabilities connected to indirect prompt injection attacks. The main conclusion is: LLMs are gullible, nobody knows how to make them completely robust. Therefore, no AI system is safe.\nIt is, of course, completely true, and the attention those attacks attract is most definitely welcome, but the main question is left unanswered: what is to be done about it?\nPurists would push for banning all external inputs that could lead to prompt injection, but, frankly, I wouldn\u0026rsquo;t be so rigid. A lot of genuinely useful applications do rely on external input, and so, I am afraid, we have to retreat to the last resort: engineering discipline.\nThere are various schemes that use a clever interaction of different models to minimize their exposure to attacks (see, for example, the CaMeL paper). I believe, though, that we don\u0026rsquo;t pay enough attention to simple input sanitization.\nA lot of such attacks rely on text hidden from the human, but visible to the machine. Detecting and removing such text is relatively trivial (in a programmatic sense), although it will require working not on the text but on the container level. This technique (called Content Disarm \u0026amp; Reconstruction, by the way) has been around for quite some time. Honestly, I am surprised that it is not implemented everywhere. It wouldn\u0026rsquo;t prevent all attacks, but it would make an unsophisticated attacker\u0026rsquo;s life harder.\nAnd to those who insist that 99% in security is a failing score, I would like to remind them of a fundamental security principle: no system is 100% secure. The role of security is to make an attack more costly than the potential benefit gained from it (see the Gordon–Loeb model).\nWith this paradigm in mind, we can build reliable and secure AI systems, even with insecure individual components.\n","permalink":"https://meshrefine.com/en/microposts/2025-08-29-documents-sanitization/","summary":"\u003cp\u003eLately, there have been numerous alerts about security vulnerabilities connected to indirect prompt injection attacks. The main conclusion is: LLMs are gullible, nobody knows how to make them completely robust. Therefore, no AI system is safe.\u003c/p\u003e\n\u003cp\u003eIt is, of course, completely true, and the attention those attacks attract is most definitely welcome, but the main question is left unanswered: what is to be done about it?\u003c/p\u003e\n\u003cp\u003ePurists would push for banning all external inputs that could lead to prompt injection, but, frankly, I wouldn\u0026rsquo;t be so rigid. A lot of genuinely useful applications do rely on external input, and so, I am afraid, we have to retreat to the last resort: engineering discipline.\u003c/p\u003e","title":"Clean your docs first"},{"content":"Some of the things you need to know about the latest GPT-5 release that evangelists don\u0026rsquo;t talk about:\nGPT-5 is not one model. What they call GPT-5 outside of API context is a router that sends your request to a model that it thinks would work most efficiently on it. You need to look at the OpenAI\u0026rsquo;s promise to provide access to everyone in this light. They provide access to the router, and you don\u0026rsquo;t know the specific configuration it applies to you and whether you will actually be able to test the most powerful model. That would undoubtedly cause completely different experiences for different users.\nGPT-5 is not a PhD. It is a pretty capable model (at least, one of them) that excels at some tasks. You can expect improvements in:\ncoding capabilities (they are impressive according to some vibe tests, but according to the SWE-bench Verified benchmark, it has just a minor lead over Claude Opus 4.1); tool calling capabilities, which are the most important for agentic workloads; other tasks where OpenAI had access to immediate feedback. While those are important improvements for a lot of areas, they don\u0026rsquo;t make the model PhD-level. The hallucination about the airfoil during the demo perfectly demonstrates that it still internalizes the most common belief, not the most current. It is a very hard problem to solve and actually a major roadblock on the way to AGI.\nDiminished hallucinations are a double-edged sword. On the one hand, the less a model hallucinates, the better, as you can trust it more. On the other hand, the more you trust the model, the more likely you are to miss actual hallucinations. In the real world, the model that never hallucinates is the best. The model that hallucinates in 0.01% of cases can be more dangerous than one that hallucinates in 10% of cases.\nMy personal impression so far is that it still has the same issue as the previous models from OpenAI, namely that it is really superficial without careful prompting. It provides you with the most shallow analysis it can get away with and hides this fact by using the very well-structured responses.\n","permalink":"https://meshrefine.com/en/microposts/2025-08-08-about-gpt-5/","summary":"\u003cp\u003eSome of the things you need to know about the latest GPT-5 release that evangelists don\u0026rsquo;t talk about:\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eGPT-5 is not one model\u003c/strong\u003e. What they call GPT-5 outside of API context is a router that sends your request to a model that it thinks would work most efficiently on it. You need to look at the OpenAI\u0026rsquo;s promise to provide access to everyone in this light. They provide access to the router, and you don\u0026rsquo;t know the specific configuration it applies to you and whether you will actually be able to test the most powerful model. That would undoubtedly cause completely different experiences for different users.\u003c/p\u003e","title":"2025-08-08"},{"content":"Google Opal\nGoogle has started a public preview for its new tool for graphical creation of multi-step AI workflows. While not a tool for production use, it is a great helper for building personal tools (and everybody should build personal tools, really, it is the biggest differentiator now).\nIt works on the Gemini platform with Gemini models (well, of course), and since it is a preview tool, it is free for now. In the future it will most likely use the Gemini Plan if the user has it. I am not sure about custom API keys, as this looks like a general public tool, but we\u0026rsquo;ll see.\nRight now it is available only in the US.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-31-google-opal/","summary":"\u003cp\u003e\u003ca href=\"https://developers.googleblog.com/en/introducing-opal/\"\u003eGoogle Opal\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003e\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n  \n  \n  \n    \n    \n    \n    \n    \n    \n    \n  \u003cpicture\u003e\n    \u003csource srcset=\"/en/microposts/2025-07-31-google-opal/google-opal_hu_df8c6253ecad9107.webp\" type=\"image/webp\" /\u003e\n  \u003cimg class=\"img-fluid\" src=\"/en/microposts/2025-07-31-google-opal/google-opal.f5b0b4e4310d106f8f86ef27ed3e0029.jpg\" alt=\"Google Opal\" loading=\"lazy\" height=\"723\" width=\"1374\" /\u003e\n\u003c/picture\u003e\n\u003c/p\u003e\n\u003cp\u003eGoogle has started a public preview for its new tool for graphical creation of multi-step AI workflows. While not a tool for production use, it is a great helper for building personal tools (and everybody should build personal tools, really, it is the biggest differentiator now).\u003c/p\u003e\n\u003cp\u003eIt works on the Gemini platform with Gemini models (well, of course), and since it is a preview tool, it is free for now. In the future it will most likely use the Gemini Plan if the user has it. I am not sure about custom API keys, as this looks like a general public tool, but we\u0026rsquo;ll see.\u003c/p\u003e","title":"2025-07-31"},{"content":"X is buzzing with a new Horizon Alpha model that is beating all previous models on various vibe tests (read: unicorns on bicycles and so on) singlehandedly. This model has 256k context window, which is a solid, albeit not the most impressive, number.\nMost probably it is a new OpenAI model (GPT-5?), as they already did the same trick before. You can try it on openrouter.ai completely for free, but remember not to provide it with any private data, as it is collected and used for model improvement.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-31-horizon-alpha-model/","summary":"\u003cp\u003e\u003ca href=\"https://x.com/OpenRouterAI/status/1950713168193282078\"\u003eX is buzzing\u003c/a\u003e with a new Horizon Alpha model that is beating all previous models on various vibe tests (read: unicorns on bicycles and so on) singlehandedly. This model has 256k context window, which is a solid, albeit not the most impressive, number.\u003c/p\u003e\n\u003cp\u003eMost probably it is a new OpenAI model (GPT-5?), as they already did the same trick before. You can try it on \u003ca href=\"https://openrouter.ai/openrouter/horizon-alpha\"\u003eopenrouter.ai\u003c/a\u003e completely for free, but remember not to provide it with any private data, as it is collected and used for model improvement.\u003c/p\u003e","title":"2025-07-31"},{"content":"Claude Code Sub Agent\nAnthropic has added the ability to create and use specialized sub-agents in Claude Code. These sub-agents use a separate context window, which allows you to run separate tasks without polluting the main context, limiting the context rot effect. You can run the created sub-agents manually or let Claude Code decide when to use them.\nWhat can be a good sub-agent? Anything that needs to be an expert in its area, doesn\u0026rsquo;t need to share the context with the main agent, can have dedicated tools, and can be designed to run self-sufficient tasks. To give you a taste of what can be made into a sub-agent, here are a couple of examples:\nGit sub-agent to which you can offload various git-related operations. As it can use just the local git state, and doesn\u0026rsquo;t need to have access to the global context, it is a good choice for a tool with its own context window. A book-writing sub-agent (I know, but I create book-like documents in a specific style for self-education). You provide it with a high-level plan, a set of materials, and one or two sections as examples, and let it work on a section separately from other sub-agents. The second example would really benefit from the ability to run several sub-agents in parallel, but it looks like I\u0026rsquo;m asking too much.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-28-claude-code-subagent/","summary":"\u003cp\u003e\u003ca href=\"https://docs.anthropic.com/en/docs/claude-code/sub-agents\"\u003eClaude Code Sub Agent\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eAnthropic has added the ability to create and use specialized sub-agents in Claude Code. These sub-agents use a separate context window, which allows you to run separate tasks without polluting the main context, limiting the context rot effect. You can run the created sub-agents manually or let Claude Code decide when to use them.\u003c/p\u003e\n\u003cp\u003eWhat can be a good sub-agent? Anything that needs to be an expert in its area, doesn\u0026rsquo;t need to share the context with the main agent, can have dedicated tools, and can be designed to run self-sufficient tasks. To give you a taste of what can be made into a sub-agent, here are a couple of examples:\u003c/p\u003e","title":"2025-07-28"},{"content":"Quoting Arvind Narayanan:\nIf we compared AI capabilities against humans with no access to tools, such as the internet, we would probably find that AI already outperformed humans at many or most cognitive tasks we perform at work. But of course this is not a helpful comparison and doesn’t tell us much about AI’s economic impacts. We are nothing without our tools.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-21-quot1/","summary":"\u003cp\u003e\u003ca href=\"https://x.com/random_walker/status/1946180439045018046\"\u003eQuoting Arvind Narayanan\u003c/a\u003e:\u003c/p\u003e\n\u003cblockquote\u003e\n\u003cp\u003eIf we compared AI capabilities against humans with no access to tools, such as the internet, we would probably find that AI already outperformed humans at many or most cognitive tasks we perform at work. But of course this is not a helpful comparison and doesn’t tell us much about AI’s economic impacts. \u003cstrong\u003eWe are nothing without our tools\u003c/strong\u003e.\u003c/p\u003e\n\u003c/blockquote\u003e","title":"2025-07-21"},{"content":"TwelveLabs video understanding models are now available in Amazon Bedrock\nAWS adds native video embeddings and video understanding models to Amazon Bedrock. It opens a lot of potential use cases for which I previously reached for Gemini models. One example of such a case is an educational system that watches how the learner performs the task and provides feedback based on the educational materials.\nBedrock had workflows to do video understanding, but it was exactly that: workflows, not native models. You can imagine what they looked like\u0026mdash;take a video, split to frames, feed frames to VLM, try to maintain temporal consistency, despair, come to terms with the system\u0026rsquo;s performance, and go on vacation.\nNow, however, there are not one, but two different native video models:\nTwelveLabs Marengo, for creating video embeddings; TwelveLabs Pegasus, for video-based text generation. Pricing of the models depends on whether your video has an audio track or not, but you should expect $2.5-$3/hour of video for Marengo and $1.8/hour for Pegasus.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-20-twelvelabs-models-in-amazon-bedrock/","summary":"\u003cp\u003e\u003ca href=\"https://aws.amazon.com/blogs/aws/twelvelabs-video-understanding-models-are-now-available-in-amazon-bedrock/\"\u003eTwelveLabs video understanding models are now available in Amazon Bedrock\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eAWS adds native video embeddings and video understanding models to Amazon Bedrock. It opens a lot of potential use cases for which I previously reached for Gemini models. One example of such a case is an educational system that watches how the learner performs the task and provides feedback based on the educational materials.\u003c/p\u003e\n\u003cp\u003eBedrock had workflows to do video understanding, but it was exactly that: workflows, not native models. You can imagine what they looked like\u0026mdash;take a video, split to frames, feed frames to VLM, try to maintain temporal consistency, despair, come to terms with the system\u0026rsquo;s performance, and go on vacation.\u003c/p\u003e","title":"2025-07-20"},{"content":"One form of context rot is what I call self-reinforced structure. When you accept a long-form model answer, you signal that this structure is acceptable, and so it tries to generate subsequent responses in a similar way. It can be destructive for any long-form creative work.\nThe only real defense is ensuring that the history the model receives doesn\u0026rsquo;t contain such replies. So it should be either prevented early or later fixed by providing a summary of the previous conversation instead of the actual history.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-20-context-rot-self-reinforcing-structure/","summary":"\u003cp\u003eOne form of \u003ca href=\"https://simonwillison.net/2025/Jun/18/context-rot/\"\u003econtext rot\u003c/a\u003e is what I call self-reinforced structure. When you accept a long-form model answer, you signal that this structure is acceptable, and so it tries to generate subsequent responses in a similar way. It can be destructive for any long-form creative work.\u003c/p\u003e\n\u003cp\u003eThe only real defense is ensuring that the history the model receives doesn\u0026rsquo;t contain such replies. So it should be either prevented early or later fixed by providing a summary of the previous conversation instead of the actual history.\u003c/p\u003e","title":"2025-07-20"},{"content":"Introducing ChatGPT agent: bridging research and action\nOpenAI released an agent that can control your own computer. It uses its advanced reasoning capabilities to plan and solve tasks in applications like Excel and PowerPoint.\nWhile Sam Altman \u0026ldquo;feels the agi\u0026rdquo; looking at how the system works, I find it incredibly clunky. Instead of concentrating on providing the models with native tools (MCP is a good step forward, although not without its problems), they try to emulate hands and eyes for them, so models can do the same things we do, but slowly and awkwardly.\nSo I would consider this type of agent a temporary workaround until we develop better machine-to-machine communication mechanisms. After this, it will be used to serve an increasingly long tail of legacy systems that will not have such machine-usable interfaces.\nP.S. Gemini mentioned a point of view I didn\u0026rsquo;t consider, namely that such systems can collect data necessary to train better embodied intelligence, meaning one that can act in the real world. It\u0026rsquo;s a perfectly valid point that shouldn\u0026rsquo;t be left without attention.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-17-chatgpt-agent/","summary":"\u003cp\u003e\u003ca href=\"https://openai.com/index/introducing-chatgpt-agent/\"\u003eIntroducing ChatGPT agent: bridging research and action\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eOpenAI released an agent that can control your own computer. It uses its advanced reasoning capabilities to plan and solve tasks in applications like Excel and PowerPoint.\u003c/p\u003e\n\u003cp\u003eWhile Sam Altman \u0026ldquo;feels the agi\u0026rdquo; looking at how the system works, I find it incredibly clunky. Instead of concentrating on providing the models with native tools (MCP is a good step forward, although not without its problems), they try to emulate hands and eyes for them, so models can do the same things we do, but slowly and awkwardly.\u003c/p\u003e","title":"2025-07-17"},{"content":"Stanford\u0026rsquo;s 2025 AI Index Report\nStanford published its annual report. It\u0026rsquo;s pretty important, because it separates speculation from pure numbers. Along with some obvious things (AI is getting better, cheaper, widespread, duh), there are some very interesting facts:\nWhile almost every organization is using AI now (78% in 2024, although no doubt, for most of them it boils down to using chatbots to compose emails), the actual results are somewhat modest. The productivity increase is on the scale of 10% (to be honest, such an increase in one year is kinda unprecedented), but the increase in revenue for most industries is just about 5%. Why? Because as with any general purpose technology, realization of full benefit would require complete rebuilding the organizational structures and processes. The problem is that no one knows how these new processes would look like, and we will have to learn from our own mistakes.\nMaybe old news, but AI provides more leverage to less experienced employees. The great equalizer of modern times. Again, that means that we need to reformulate our approach to team staffing. I would only add that it can help only if you have some remote understanding of what you\u0026rsquo;re doing, so those who apply for entry positions, do your homework well.\nThe number of AI-related incidents continues to rise. We see a twofold increase in 2024 vs 2023, and this is before frantic adoption of Agents and MCPs we see in 2025. So, we need to brace ourselves and be ready for more and more data leaks and integrity breaches.\nThe report contains a lot more nuggets, but it\u0026rsquo;s almost 500 pages long, so I would really recommend to use AI to extract what you fancy.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-17-stanfordhai2025/","summary":"\u003cp\u003e\u003ca href=\"https://hai.stanford.edu/ai-index/2025-ai-index-report\"\u003eStanford\u0026rsquo;s 2025 AI Index Report\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eStanford published its annual report. It\u0026rsquo;s pretty important, because it separates speculation from pure numbers. Along with some obvious things (AI is getting better, cheaper, widespread, duh), there are some very interesting facts:\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003e\n\u003cp\u003eWhile almost every organization is using AI now (78% in 2024, although no doubt, for most of them it boils down to using chatbots to compose emails), the actual results are somewhat modest.  The productivity increase is on the scale of 10% (to be honest, such an increase in one year is kinda unprecedented), but the increase in revenue for most industries is just about 5%. Why? Because as with any general purpose technology, realization of full benefit would require complete rebuilding the organizational structures and processes. The problem is that no one knows how these new processes would look like, and we will have to learn from our own mistakes.\u003c/p\u003e","title":"2025-07-17"},{"content":"Voxtral\nMistral introduces Voxtral, a family of open-source speech recognition and understanding models. It\u0026rsquo;s about time. We haven\u0026rsquo;t seen a comparable open-source model since OpenAI\u0026rsquo;s Whisper, and that was quite a while ago.\nThe models are provided in 3B and 24B sizes and outperform Whisper on most benchmarks. However, they require more powerful hardware, as the largest Whisper variant is just 1.5B. This is a direct consequence of it also being a regular language model. Another consequence is that controlling them in a pure transcription setting would be harder.\nThe models are available on Hugging Face as well as through the Mistral API and their LeChat.\nWhat they also currently lack is diarization (speaker recognition) support. It\u0026rsquo;s on the roadmap, but in the meantime, we still have to use somewhat clunky pyannote-audio for this purpose.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-17-voxtral/","summary":"\u003cp\u003e\u003ca href=\"https://mistral.ai/news/voxtral\"\u003eVoxtral\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eMistral introduces Voxtral, a family of open-source speech recognition and understanding models. It\u0026rsquo;s about time. We haven\u0026rsquo;t seen a comparable open-source model since OpenAI\u0026rsquo;s Whisper, and that was quite a while ago.\u003c/p\u003e\n\u003cp\u003eThe models are provided in 3B and 24B sizes and outperform Whisper on most benchmarks. However, they require more powerful hardware, as the largest Whisper variant is just 1.5B. This is a direct consequence of it also being a regular language model. Another consequence is that controlling them in a pure transcription setting would be harder.\u003c/p\u003e","title":"2025-07-17"},{"content":"Strategic Intelligence in Large Language Models: Evidence from evolutionary Game Theory\nThe paper shows that different models behave completely differently when placed in game theory settings. What that means is that testing and evals are playing an increasingly critical role in developing agentic systems, as updating or changing the underlying model will lead to unpredictable changes in an agent\u0026rsquo;s behaviour.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-17-ai-prisoners-dilemma/","summary":"\u003cp\u003e\u003ca href=\"https://arxiv.org/abs/2507.02618\"\u003eStrategic Intelligence in Large Language Models: Evidence from evolutionary Game Theory\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eThe paper shows that different models behave completely differently when placed in game theory settings. What that means is that testing and evals are playing an increasingly critical role in developing agentic systems, as updating or changing the underlying model will lead to unpredictable changes in an agent\u0026rsquo;s behaviour.\u003c/p\u003e","title":"2025-07-17"},{"content":"Introducing Kiro\nAWS jumps into the agentic IDEs bandwagon with Kiro. To separate itself from vibe-coding approach, which accumulated a considerable amount of ill repute, they emphasize the \u0026ldquo;spec-driven development\u0026rdquo; method. That means that the agent first helps the user to create a full requirements document for the feature, then it analyzes the existing code base, and only after that it starts implementing.\nThis approach definitely makes sense, and it\u0026rsquo;s a step forward from blindly running into the fray that is vibe-coding. The fact that those specs are updating along with the code changes makes them even more valuable, minimizing the problem of stale documentation. Hooks can run repeated agentic tasks, such as making sure the new feature has sufficient tests, automatically.\nIt is interesting to watch how different tools adopt different methodologies, as it allows the developers community to find and disseminate the techniques that really work.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-15-aws-introduces-kiro/","summary":"\u003cp\u003e\u003ca href=\"https://kiro.dev/blog/introducing-kiro/\"\u003eIntroducing Kiro\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eAWS jumps into the agentic IDEs bandwagon with Kiro. To separate itself from vibe-coding approach, which accumulated a considerable amount of ill repute, they emphasize the \u0026ldquo;spec-driven development\u0026rdquo; method. That means that the agent first helps the user to create a full requirements document for the feature, then it analyzes the existing code base, and only after that it starts implementing.\u003c/p\u003e\n\u003cp\u003eThis approach definitely makes sense, and it\u0026rsquo;s a step forward from blindly running into the fray that is vibe-coding. The fact that those specs are updating along with the code changes makes them even more valuable, minimizing the problem of stale documentation. Hooks can run repeated agentic tasks, such as making sure the new feature has sufficient tests, automatically.\u003c/p\u003e","title":"2025-07-15"},{"content":"TIL: ccusage — a nice tool to track and analyze the Claude Code usage.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-15-0d12c26b/","summary":"\u003cp\u003eTIL: \u003ca href=\"https://simonwillison.net/2025/Jul/14/ccusage/\"\u003eccusage\u003c/a\u003e — a nice tool to track and analyze the Claude Code usage.\u003c/p\u003e","title":"2025-07-15"},{"content":"Anthropic released 4 new courses in its academy:\nClaude Code in Action with practical advice on using the CLI agent. Claude with the Anthropic API, a comprehensive course on using all current API capabilities, from single-shot text generation to agents. Introduction to Model Context Protocol and Model Context Protocol: Advanced Topics for those interested in MCP. Each course comes in video and text formats and provides a certificate of completion.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-14-6fe8db2d/","summary":"\u003cp\u003eAnthropic released 4 new courses in \u003ca href=\"https://www.anthropic.com/learn\"\u003eits academy\u003c/a\u003e:\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003e\u003ca href=\"https://anthropic.skilljar.com/claude-code-in-action\"\u003eClaude Code in Action\u003c/a\u003e with practical advice on using the CLI agent.\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"https://anthropic.skilljar.com/claude-with-the-anthropic-api\"\u003eClaude with the Anthropic API\u003c/a\u003e, a comprehensive course on using all current API capabilities, from single-shot text generation to agents.\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"https://anthropic.skilljar.com/introduction-to-model-context-protocol\"\u003eIntroduction to Model Context Protocol\u003c/a\u003e and \u003ca href=\"https://anthropic.skilljar.com/model-context-protocol-advanced-topics\"\u003eModel Context Protocol: Advanced Topics\u003c/a\u003e for those interested in MCP.\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003eEach course comes in video and text formats and provides a certificate of completion.\u003c/p\u003e","title":"2025-07-14"},{"content":"TIL: DBML - Database Markup Language\nA markup language for DB schema description that can be useful to provide it to AI tools.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-14-22cb7f0e/","summary":"\u003cp\u003eTIL: \u003ca href=\"https://dbml.dbdiagram.io/home/\"\u003eDBML - Database Markup Language\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eA markup language for DB schema description that can be useful to provide it to AI tools.\u003c/p\u003e","title":"2025-07-14"},{"content":"Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity\nAn important and well-conducted study showing that modern AI-coding tools are detrimental to the developers who:\nare well-experienced; work on large code bases; know the code base like the back of their hand. Interestingly, although on average they closed their tasks 19% slower, they still percieved that AI agents provided about 20% speedup!\nThe authors note that this effect can diminish with new generations of Agents, and, most importantly, there is a wide selection of tasks that already show performance improvement.\nSo if you\u0026rsquo;re not deeply familiar with the code base, have limited experience with technologies used, or are developing a greenfield project, an AI agent will likely help.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-11-20dc3714/","summary":"\u003cp\u003e\u003ca href=\"https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/\"\u003eMeasuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eAn important and well-conducted study showing that modern AI-coding tools are detrimental to the developers who:\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eare well-experienced;\u003c/li\u003e\n\u003cli\u003ework on large code bases;\u003c/li\u003e\n\u003cli\u003eknow the code base like the back of their hand.\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eInterestingly, although on average they closed their tasks 19% slower, they still percieved that AI agents provided about 20% speedup!\u003c/p\u003e\n\u003cp\u003eThe authors note that this effect can diminish with new generations of Agents, and, most importantly, there is a wide selection of tasks that already show performance improvement.\u003c/p\u003e","title":"The AI Productivity Paradox"},{"content":"It is very important to understand that AI models are not deterministic and they cannot be made deterministic without severely restricting the environment they run in. Fixing seeds doesn\u0026rsquo;t help. Setting temperature to 0 doesn\u0026rsquo;t help. Every small floating point rounding error can ultimately lead to drastically different results.\nThe complex models are, in essence, chaotic systems and should be treated as such.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-09-259fafd6/","summary":"\u003cp\u003eIt is very important to understand that AI models are not deterministic and they cannot be made deterministic without severely restricting the environment they run in. Fixing seeds doesn\u0026rsquo;t help. Setting temperature to 0 doesn\u0026rsquo;t help. Every small floating point rounding error can ultimately lead to drastically different results.\u003c/p\u003e\n\u003cp\u003eThe complex models are, in essence, chaotic systems and should be treated as such.\u003c/p\u003e","title":"2025-07-09"},{"content":"AI coding agents are just another tool in a good developer\u0026rsquo;s toolbox. They start with a slab of stone, and use those agents as a metaphorical sledgehammer to give it the rough form they envision. Then, they reach for AI-assisted coding tools, such as GitHub Copilot, to work with more precision, akin to a smaller hammer. And finally, they use the smallest chisel to carve out the finest details by hand.\nHow funny it is to hear that the art of programming is dead because we no longer carve the slabs with those small chisels alone.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-07-ae052c28/","summary":"\u003cp\u003eAI coding agents are just another tool in a good developer\u0026rsquo;s toolbox. They start with a slab of stone, and use those agents as a metaphorical sledgehammer to give it the rough form they envision. Then, they reach for AI-assisted coding tools, such as GitHub Copilot, to work with more precision, akin to a smaller hammer. And finally, they use the smallest chisel to carve out the finest details by hand.\u003c/p\u003e","title":"2025-07-07"},{"content":"A funny thing: Gemini Deep Research is programmatically discouraged from finishing early. I found that out when I tried to use it to extract information from unstructured text and fill in a template. Despite its being a research tool, it has everything necessary for such a task: access to Google Docs, ability to create long-form documents and the meticulous agentic flow. Of course, if we want to just restructure the document, it is important that the model does not use internet search at all.\nThe agent did the work quite well and quickly. However, when it was about to finish, it received several \u0026ldquo;continue research\u0026rdquo; urges from the programmatic orchestrator, and guess what? It started browsing the internet for the missing information.\nThe moral is: you can be creative with such systems, but you should expect the scaffolding to throw a wrench in the works.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-05-55a52a46/","summary":"\u003cp\u003eA funny thing: Gemini Deep Research is programmatically discouraged from finishing early. I found that out when I tried to use it to extract information from unstructured text and fill in a template. Despite its being a research tool, it has everything necessary for such a task: access to Google Docs, ability to create long-form documents and the meticulous agentic flow.  Of course, if we want to just restructure the document, it is important that the model does not use internet search at all.\u003c/p\u003e","title":"2025-07-05"},{"content":"After Weeks of Markey Raising the Alarm, Senate Strikes AI Moratorium from Budget Reconciliation Bill Overnight in Overwhelming 99-1 Vote\nSo, no moratorium on state-level AI regulations. Those developing AI systems, get ready to learn and implement compliance measures for 50 different states. Woe to us.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-02-72f7e44e/","summary":"\u003cp\u003e\u003ca href=\"https://www.markey.senate.gov/news/press-releases/after-weeks-of-markey-raising-the-alarm-senate-strikes-ai-moratorium-from-budget-reconciliation-bill-overnight-in-overwhelming-99-1-vote\"\u003eAfter Weeks of Markey Raising the Alarm, Senate Strikes AI Moratorium from Budget Reconciliation Bill Overnight in Overwhelming 99-1 Vote\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eSo, no moratorium on state-level AI regulations. Those developing AI systems, get ready to learn and implement compliance measures for 50 different states. Woe to us.\u003c/p\u003e","title":"2025-07-02"},{"content":"Cloudflare Just Changed How AI Crawlers Scrape the Internet-at-Large; Permission-Based Approach Makes Way for A New Business Model\nCloudflare makes a large step towards data monetization for AI training. Now all new data hosted on Cloudflare is inaccessible to AI crawlers by default. There is also a new option for the page to return code HTTP 402 (\u0026ldquo;Payment required\u0026rdquo;) and charge for access. Of course, Cloudflare will be an intermediary, which gives it significant control and financial power.\nMoreover, it could spark an era of \u0026ldquo;dark crawling\u0026rdquo;, which will undoubtedly lead to cat-and-mouse games of crawer detection.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-02-c80a82b1/","summary":"\u003cp\u003e\u003ca href=\"https://www.cloudflare.com/ru-ru/press-releases/2025/cloudflare-just-changed-how-ai-crawlers-scrape-the-internet-at-large/\"\u003eCloudflare Just Changed How AI Crawlers Scrape the Internet-at-Large; Permission-Based Approach Makes Way for A New Business Model\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eCloudflare makes a large step towards data monetization for AI training. Now all new data hosted on Cloudflare is inaccessible to AI crawlers by default. There is also a new option for the page to return code \u003ca href=\"https://blog.cloudflare.com/introducing-pay-per-crawl/\"\u003eHTTP 402 (\u0026ldquo;Payment required\u0026rdquo;)\u003c/a\u003e and charge for access. Of course, Cloudflare will be an intermediary, which gives it significant control and financial power.\u003c/p\u003e","title":"2025-07-02"},{"content":"AI-assisted coding for teams that can\u0026rsquo;t get away with vibes\nA very nice guide to using AI agents for coding efficiently. It\u0026rsquo;s brief, but it gives the idea where to dig.\n","permalink":"https://meshrefine.com/en/microposts/2025-07-01-agentic-coding-advice/","summary":"\u003cp\u003e\u003ca href=\"https://blog.nilenso.com/blog/2025/05/29/ai-assisted-coding/\"\u003eAI-assisted coding for teams that can\u0026rsquo;t get away with vibes\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eA very nice guide to using AI agents for coding efficiently. It\u0026rsquo;s brief, but it gives the idea where to dig.\u003c/p\u003e","title":"2025-07-01"},{"content":"Google search head reveals \u0026ldquo;game system\u0026rdquo; approach as AI mode begins advertising rollout\nGoogle will show ads in its AI mode. An entirely expected move from the search giant, given that AI search and other features break the traditional ads model. It will be interesting to see how those new ad integration techniques will evolve, and what the new generation of ad blockers will look like.\n","permalink":"https://meshrefine.com/en/microposts/2025-06-30-google-to-roll-out-ads-in-ai-mode/","summary":"\u003cp\u003e\u003ca href=\"https://ppc.land/google-search-head-reveals-game-system-approach-as-ai-mode-begins-advertising-rollout/\"\u003eGoogle search head reveals \u0026ldquo;game system\u0026rdquo; approach as AI mode begins advertising rollout\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eGoogle will show ads in its AI mode. An entirely expected move from the search giant, given that AI search and other features break the traditional ads model. It will be interesting to see how those new ad integration techniques will evolve, and what the new generation of ad blockers will look like.\u003c/p\u003e","title":"2025-06-30"},{"content":"How Long Contexts Fail\nHow to Fix Your Context\nLong context is your friend\u0026hellip; when we are talking about summarization and retrieval. For agentic workflow it is often detrimental due to reasons such as:\nContext poisoning, where the model hallucinates and messes with its own context. The ripples are powerful and they die slowly;\nContext distraction, where the model starts to repeat itself instead of trying new strategies;\nContext confusion, which happens when one gives the model too many tools (sometimes 2 is too many);\nContext clash, where the model cannot let go of its initial wrong assumptions. And it assumes something each turn of the conversation, be it a tool call, file read, or user\u0026rsquo;s request.\nThose problems become especially emphasized when the context starts getting too long.\nFortunately, there are some methods to minimize the impact of those issues:\nRAG. Rumors about its death are greatly exaggerated. It is still a very powerful tool to control the context length;\nTool loadout. Give the model only the tools it needs. A task changed? Select another set of tools. Or let a RAG system or another LLM do that.\nContext quarantine. You can delegate, then why the agent shouldn\u0026rsquo;t be able to? Split tasks into subtasks with their own contexts, consume the results.\nContext pruning. Trim the context mercilessly. The model doesn\u0026rsquo;t need to remember its reasoning traces or tool calls results from 10 turns back. And keeping everything in some structured format lets you efficiently develop strategies for pruning.\nContext summarization. When context gets too long, summarize it and throw away. Don\u0026rsquo;t forget to reinitialize the agent after that. If you have files with instructions, add them to the resulting summary.\nContext offloading. Give the model a scratchpad. Or, better yet, several. Let it write down, forget, consult, repeat.\nAll of that will make the agent system not only more reliable, but also cheaper.\n","permalink":"https://meshrefine.com/en/posts/long-context-failures-and-fixes/","summary":"\u003cp\u003e\u003ca href=\"https://www.dbreunig.com/2025/06/22/how-contexts-fail-and-how-to-fix-them.html\"\u003eHow Long Contexts Fail\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003e\u003ca href=\"https://www.dbreunig.com/2025/06/26/how-to-fix-your-context.html\"\u003eHow to Fix Your Context\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eLong context is your friend\u0026hellip; when we are talking about summarization and retrieval. For agentic workflow it is often detrimental due to reasons such as:\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eContext poisoning\u003c/strong\u003e, where the model hallucinates and messes with its own context. The ripples are powerful and they die slowly;\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eContext distraction\u003c/strong\u003e, where the model starts to repeat itself instead of trying new strategies;\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eContext confusion\u003c/strong\u003e, which happens when one gives the model too many tools (sometimes 2 is too many);\u003c/p\u003e","title":"The Treachery of Memory: On Long Contexts and Agentic Failures"},{"content":"When you work with an AI coding agent, you really want to provide the agent with a lot of instructions related to workflows, codebase organization, styling guidelines and more. Those instructions should be reusable and composable, so you could have uniform codebases everywhere. The repo provides examples of such instructions for web development and Windsurf.\nI would recommend creating a similar set of files for the projects you work on, maintain them, and compile them into AGENTS.md/CLAUDE.md/GEMINI.md using a simple cp command:\n$ cp Workflow.md Styles.md Structure.md \u0026gt; AGENTS.md ","permalink":"https://meshrefine.com/en/microposts/2025-06-28-b5dadddc/","summary":"\u003cp\u003eWhen you work with an AI coding agent, you \u003cem\u003ereally\u003c/em\u003e want to provide the agent with a lot of instructions related to workflows, codebase organization, styling guidelines and more. Those instructions should be reusable and composable, so you could have uniform codebases everywhere. The \u003ca href=\"https://github.com/Shamail/ai-coding-template\"\u003erepo\u003c/a\u003e provides examples of such instructions for web development and Windsurf.\u003c/p\u003e\n\u003cp\u003eI would recommend creating a similar set of files for the projects you work on, maintain them, and compile them into \u003ccode\u003eAGENTS.md\u003c/code\u003e/\u003ccode\u003eCLAUDE.md\u003c/code\u003e/\u003ccode\u003eGEMINI.md\u003c/code\u003e using a simple \u003ccode\u003ecp\u003c/code\u003e command:\u003c/p\u003e","title":"AI-Assisted Coding Template"},{"content":"Microsoft\u0026rsquo;s rekindling of Three Mile Island nuclear plant is ahead of schedule\nDemand for AI is resurrecting nuclear power—the most efficient and carbon-free power source we currently posess. This serves as a stark reminder that one cannot consider immediate effects in isolation from their second- and third-order consequences.\n","permalink":"https://meshrefine.com/en/microposts/2025-06-28-b46c7630/","summary":"\u003cp\u003e\u003ca href=\"https://www.techradar.com/pro/microsofts-rekindling-of-three-mile-island-nuclear-plant-is-ahead-of-schedule\"\u003eMicrosoft\u0026rsquo;s rekindling of Three Mile Island nuclear plant is ahead of schedule\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eDemand for AI is resurrecting nuclear power—the most efficient and carbon-free power source we currently posess. This serves as a stark reminder that one cannot consider immediate effects in isolation from their second- and third-order consequences.\u003c/p\u003e","title":"2025-06-28"},{"content":"State of Cybersecurity Resilience 2025\nOnly one in ten organizations is sufficiently protected against AI-related threats. Hasty adoption of AI-based tools dramatically increases the attack surface. It is crucial to carefully and methodically develop the security strategy for each system, collect the best practices and educate the stakeholders.\n","permalink":"https://meshrefine.com/en/microposts/2025-06-28-3169d51b/","summary":"\u003cp\u003e\u003ca href=\"https://www.accenture.com/us-en/insights/security/state-cybersecurity-2025\"\u003eState of Cybersecurity Resilience 2025\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eOnly one in ten organizations is sufficiently protected against AI-related threats. Hasty adoption of AI-based tools dramatically increases the attack surface. It is crucial to carefully and methodically develop the security strategy for each system, collect the best practices and educate the stakeholders.\u003c/p\u003e","title":"2025-06-28"},{"content":"Swift Android Workgroup\nApple formed a workgroup to support Android app development using Swift. It can be interesting for companies that already have iOS applications and would like to reuse the business logic on Android. The UI and other platform-specific components, however, still have to be written using native frameworks and libraries.\n","permalink":"https://meshrefine.com/en/microposts/2025-06-27-cb648728/","summary":"\u003cp\u003e\u003ca href=\"https://www.swift.org/android-workgroup/\"\u003eSwift Android Workgroup\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eApple formed a workgroup to support Android app development using Swift.\nIt can be interesting for companies that already have iOS applications and would like to reuse the business logic on Android.  The UI  and other platform-specific components, however, still have to be written using native frameworks and libraries.\u003c/p\u003e","title":"2025-06-27"},{"content":"Build and share AI-powered apps with Claude\nA new feature from Anthropic that allows you to create and share small AI-powered applications using artifacts. These apps use the end user\u0026rsquo;s Claude account for that. Turns out that Gemini already has had this feature for some time.\n","permalink":"https://meshrefine.com/en/microposts/2025-06-26-93a301c0/","summary":"\u003cp\u003e\u003ca href=\"https://www.anthropic.com/news/claude-powered-artifacts\"\u003eBuild and share AI-powered apps with Claude\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eA new feature from Anthropic that allows you to create and share small AI-powered applications using artifacts. These apps use the end user\u0026rsquo;s Claude account for that.\n\u003ca href=\"https://x.com/emollick/status/1938091740121935929\"\u003eTurns out\u003c/a\u003e that Gemini already has had this feature for some time.\u003c/p\u003e","title":"2025-06-26"},{"content":"Gemini CLI: your open-source AI agent\nNow Google has its own CLI coding agent. Massive context, generous free tier (it\u0026rsquo;s free for most cases), and fully open-source. I\u0026rsquo;m yet to try it, but it looks impressive.\n","permalink":"https://meshrefine.com/en/microposts/2025-06-25-9bb932fa/","summary":"\u003cp\u003e\u003ca href=\"https://blog.google/technology/developers/introducing-gemini-cli-open-source-ai-agent/\"\u003eGemini CLI: your open-source AI agent\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eNow Google has its own CLI coding agent. Massive context, generous free tier (it\u0026rsquo;s free for most cases), and fully open-source. I\u0026rsquo;m yet to try it, but it looks impressive.\u003c/p\u003e","title":"2025-06-25"},{"content":"Anthropic wins ruling on AI training in copyright lawsuit but must face trial on pirated books\nA very important copyright legal precedent. TL;DR: Training models on legally acquired copyrighted materials is fair use. The model\u0026rsquo;s creators need to make sure that its output is \u0026ldquo;quintessentially transformative\u0026rdquo;.\nThe illegal acquisition of training data is punishable as before.\n","permalink":"https://meshrefine.com/en/microposts/2025-06-25-f08b90df/","summary":"\u003cp\u003e\u003ca href=\"https://apnews.com/article/anthropic-ai-fair-use-copyright-pirated-libraries-1e5cece51c2e4bd0bb21d94de2abb035\"\u003eAnthropic wins ruling on AI training in copyright lawsuit but must face trial on pirated books\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eA very important copyright legal precedent. TL;DR: Training models on legally acquired copyrighted materials is fair use. The model\u0026rsquo;s creators need to make sure that its output is \u0026ldquo;quintessentially transformative\u0026rdquo;.\u003c/p\u003e\n\u003cp\u003eThe illegal acquisition of training data is punishable as before.\u003c/p\u003e","title":"2025-06-25"},{"content":"I usually approach shiny new things with a healthy dose of skepticism. Until recently, this was precisely my attitude toward multi-agent systems. This is hardly surprising, given the immense hype surrounding them and the conspicuous absence of genuinely successful examples. Most implementations that actually worked fell into one of the following categories:\nAgentic systems following a predefined plan. These are essentially LLMs with tools, trained to automate a very specific process. This approach allows each step to be tested individually and its results verified. Such systems are typically described as a directed acyclic graph (DAG), sometimes dynamic, and developed using now-standard primitives from frameworks like LangChain and Griptape1. The early implementation of Gemini Deep Research operated this way: first, a search plan was created, then the search was executed, and finally, the results were compiled. Solutions operating in systems with a feedback loop. Various Claude Code, Cursor, and other code-generating agents fall into this group. The stronger the feedback loop—that is, the better the tooling and the stricter the type checking—the greater the chance they won\u0026rsquo;t completely wreck your codebase2. Models trained using Reinforcement Learning, such as those with interleaved thinking, like OpenAI\u0026rsquo;s o3. This is a separate, very interesting conversation, but even these models have a certain modus operandi defined by the specifics of their training. Meanwhile, open-ended multi-agent systems have largely remained in the proof-of-concept stage due to their general unreliability. The community lacked a clear understanding of where and how to implement them. This was the case until Anthropic published a deeply technical article on how they developed their Deep Research system. It defined a reasonably clear framework for building such systems, and that is what we will examine today.\nThe Core Idea The most important contribution of this article is the identification of a design pattern for multi-agent systems with dynamic orchestration. Yes, I\u0026rsquo;m drawing a direct analogy to the classic design patterns from the world of programming.\nClassic patterns are densely packed nuggets of wisdom from architects and programmers who have built hundreds of thousands of software systems. By analyzing their work, they identified certain regularities, which they formalized to raise the level of abstraction for architectural problems and to facilitate communication.\nA good pattern consists of3:\nA catchy name. This is an essential component; without it, the pattern simply won\u0026rsquo;t stick. A description of the problem it solves. Typically, the problem is general enough to warrant abstraction. A description of the pattern itself. And a description of when it should not be applied. Now let\u0026rsquo;s look at what the Anthropic engineers present in their article:\nThe Name The article calls it the \u0026ldquo;orchestrator-worker\u0026rdquo; pattern, which captures the essence but loses a key distinction from the classic pattern: the dynamic nature and adaptation of tasks for the workers based on the initial problem. I believe this is a significant enough feature to be reflected in the name. Other names they use—Advanced Research, multi-agent research system—are more about describing the application domain. Therefore, I will henceforth call it the \u0026ldquo;Adaptive Orchestrator,\u0026rdquo; or AdOrc4.\nThe Problem Description This unpredictability makes AI agents particularly well-suited for research tasks. Research demands the flexibility to pivot or explore tangential connections as the investigation unfolds. The model must operate autonomously for many turns, making decisions about which directions to pursue based on intermediate findings. A linear, one-shot pipeline cannot handle these tasks.\nThe essence of search is compression: distilling insights from a vast corpus. Subagents facilitate compression by operating in parallel with their own context windows, exploring different aspects of the question simultaneously before condensing the most important tokens for the lead research agent. Each subagent also provides separation of concerns—distinct tools, prompts, and exploration trajectories—which reduces path dependency and enables thorough, independent investigations.\n…\nOur internal evaluations show that multi-agent research systems excel especially for breadth-first queries that involve pursuing multiple independent directions simultaneously.\nHere, the engineers clearly indicate where the pattern performs well:\nIn cases where the plan needs to be modified based on intermediate results. These are not deterministic business processes; they are explorations of the surrounding world. Search and research tasks fit perfectly here. Where we hit the technical limitations of a single agent. The primary limitation is the context window, which brings along latency, high inference costs, and some quality degradation due to the nature of the attention mechanism. Finally, in situations where a large number of independent, parallel subtasks can be launched. The pattern shines in tasks like patent searches or due diligence—areas where humans also work in parallel. The Pattern Description Our Research system uses a multi-agent architecture with an orchestrator-worker pattern, where a lead agent coordinates the process while delegating to specialized subagents that operate in parallel.\nThe multi-agent architecture in action: user queries flow through a lead agent that creates specialized subagents to search for different aspects in parallel (figure from the original post by Anthropic).\nWhen a user submits a query, the lead agent analyzes it, develops a strategy, and spawns subagents to explore different aspects simultaneously. As shown in the diagram above, the subagents act as intelligent filters by iteratively using search tools to gather information, in this case on AI agent companies in 2025, and then returning a list of companies to the lead agent so it can compile a final answer.\nTraditional approaches using Retrieval Augmented Generation (RAG) use static retrieval. That is, they fetch some set of chunks that are most similar to an input query and use these chunks to generate a response. In contrast, our architecture uses a multi-step search that dynamically finds relevant information, adapts to new findings, and analyzes results to formulate high-quality answers.\nThe structure of the pattern is clear from the description:\nThe system consists of an orchestrator and workers. These are LLMs (or LMMs5) with access to tools. The orchestrator can, in the general case, assign specific tools to specific workers. The system receives a task and a description of the desired outcome. The process begins by creating a plan where tasks can be executed by the orchestrator itself or by specialized workers, provided the conditions listed above are met. At each step, the orchestrator launches the workers, which perform actions and return the resulting data. At the end of the cycle, a completion condition is checked. The process either returns to step 3, where the orchestrator modifies the plan, or it terminates and returns the result to the user. Here are a few possible termination reasons: The task requirements have been met (successful exit). The allocated budget has been exceeded. The specified number of iterations has been exceeded. Convergence (no significant improvements over the last few iterations). graph TD subgraph \u0026#34;The AdOrc Cycle\u0026#34; A(\u0026#34;1\\. Receive task and desired outcome\u0026#34;) --\u0026gt; B(\u0026#34;2\\. Build / Modify plan\u0026#34;); B --\u0026gt; C(\u0026#34;3\\. Execute workers and get results\u0026#34;); C --\u0026gt; D{\u0026#34;4\\. Check completion condition\u0026#34;}; D -- \u0026#34;No, refinement needed\u0026#34; --\u0026gt; B; D -- \u0026#34;Yes, goal achieved\u0026#34; --\u0026gt; E(\u0026#34;Return result to user\u0026#34;); end subgraph \u0026#34;Termination Reasons\u0026#34; F[\u0026#34;-Result meets requirements\u0026lt;br/\u0026gt;-Budget or iteration limit exceeded\u0026lt;br/\u0026gt;-Convergence (no recent improvements)\u0026#34;]; end D -.- F; The Limitations Now let\u0026rsquo;s look at where this pattern underperforms. The authors write:\n… in practice, these architectures burn through tokens fast. In our data, agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats. For economic viability, multi-agent systems require tasks where the value of the task is high enough to pay for the increased performance. Further, some domains that require all agents to share the same context or involve many dependencies between agents are not a good fit for multi-agent systems today. For instance, most coding tasks involve fewer truly parallelizable tasks than research, and LLM agents are not yet great at coordinating and delegating to other agents in real time.\n…\nSubagent output to a filesystem to minimize the ‘game of telephone.’ Direct subagent outputs can bypass the main coordinator for certain types of results, improving both fidelity and performance.\nAs we can see, this approach has its drawbacks:\nCost. Running such a system is expensive, so it should be used for tasks with relatively high economic value, such as analyzing case law, reviewing articles in a specific scientific field, or gathering feedback and advice when learning a new technology. Agent independence. Yes, sometimes this is a disadvantage, for instance, when agents need to maintain a large shared context. An example would be a document processing system that analyzes a document from multiple perspectives simultaneously. In the Deep Research implementation, workers sometimes had to communicate via the filesystem, which should be seen as a workaround. Latency. The cyclical nature of research using powerful (and therefore slow) models means that waiting several minutes for a result is normal. This requires a special approach to user interaction design, making it unsuitable for a vast range of applications. Consequently, AdOrc is often not suitable for scenarios like:\nWriting new code. As is well known, writing code doesn\u0026rsquo;t always play well with parallelism. Decomposing tasks so that team members don\u0026rsquo;t step on each other\u0026rsquo;s toes is a major headache. However, this pattern could be useful in other aspects of software development, such as onboarding to a new codebase, refactoring, and debugging. Automating business processes. Most are relatively well-formalized and are better solved by agents with a fixed plan. Recently, case studies of larger, more flexible automations have appeared, but they don\u0026rsquo;t provide enough detail to assess their effectiveness and reliability. Knowledge base search. While the pattern is applicable here in principle, its high latency and cost make classic RAG systems a better fit for such tasks. Creating AGI or ASI. No, I\u0026rsquo;m not saying this pattern is inapplicable to AGI. It\u0026rsquo;s just that nobody knows what is applicable. Not-So-Harmful Advice Now that we\u0026rsquo;ve covered what I consider the post\u0026rsquo;s main contribution to our technical field, let\u0026rsquo;s look at some of the advice Anthropic\u0026rsquo;s developers offer to builders of similar agentic systems. Since the developers accompanied their points with excellent notes, I won\u0026rsquo;t repeat them all, limiting myself to the ones that resonated most with me.\nStart evaluating from the very beginning, even with a small sample. Model evaluation is expensive, complex, and confusing, which is why many products settle for \u0026ldquo;vibe checking,\u0026rdquo; or as it\u0026rsquo;s also known, \u0026ldquo;I tried this prompt in ChatGPT, and it seemed to work.\u0026rdquo; This path leads nowhere (or to multi-million dollar lawsuits, product failure, or a system jailbreak—underline as appropriate). Building an effective and automated eval and red teaming system to catch problems before they surface is a critical engineering practice. Not to mention that trying to improve system prompts without a reliable way to evaluate them is like hitting a piñata blindfolded. For developing evals, you can use projects like promptfoo and DeepEval, which support many useful metrics and LLM-as-judge out of the box.\nLLM-as-judge scales if you prepare it correctly. Yes, but its preparation is a special kind of art. Different LLMs evaluate the same output in completely different ways. The article suggests using an LLM to assign scores from 0 to 1. This directly contradicts well-known research showing that even powerful models cannot consistently assign such scores. The most reliable method of using LLM-as-judge is pairwise comparison of two results, often combined with majority voting and swapping the order of options. In short, something doesn\u0026rsquo;t quite add up with this piece of advice. However, it\u0026rsquo;s entirely possible that for the Deep Research implementation, these scores worked well enough.\nHuman evaluation catches what automation misses. Human evaluation is an expensive and highly subjective process. But you cannot skip this step, because it is the only way to identify corner cases not anticipated by your tests. The article gives an example of how a tester noticed that an early version of the system was being baited by SEO and was ignoring content-rich scientific papers and personal blogs. LLM-as-judge and other metrics couldn\u0026rsquo;t detect this on their own because they didn\u0026rsquo;t consider the resource type as an input parameter. After adding this parameter and certain heuristics to the prompt, the model\u0026rsquo;s behavior improved, and the tests were adapted to account for it.\nAgents have state and accumulating error. Oh, this is the very problem that breaks multi-step agents without a feedback loop. It\u0026rsquo;s probability theory, and you can\u0026rsquo;t argue with it. If an agent has a 99% chance of completing a step correctly, what\u0026rsquo;s the probability of correctly completing a 10-step process? 90%. A 30-step process? 74%. A 100-step process will fail in 2/3 of cases. And this is an idealized situation. For a stochastic LLM operating in the messy real world, the probability of problems is significantly higher. We\u0026rsquo;re not just talking about predictable technical failures (a software defect, a weird encoding) but about issues specific to the model itself: hallucinations, logical leaps, context contamination, etc.\nWhat makes it worse is that the consequences of errors persist in the agent\u0026rsquo;s state and \u0026ldquo;poison\u0026rdquo; all subsequent steps. The solution is either to introduce intermediate feedback, which isn\u0026rsquo;t always possible, or to limit the number of turns. It is the combination of these methods that allows the AdOrc pattern, and Deep Research in particular, to function properly. The number of turns for the orchestrator is small, but at each step, it launches multiple workers that are allowed to fail without seriously impacting the final result. At the same time, it receives information about all technical failures, providing it with a feedback loop to adapt and find workarounds.\nDebugging benefits from new approaches. On this point, the post becomes disappointingly concise, even though this is precisely the information needed to build reliable agentic systems. Anthropic mentions logging decision-making patterns and interaction structures but doesn\u0026rsquo;t go into detail. However, it\u0026rsquo;s highly likely they have a fairly robust observability system in place:\nAll metadata (spans) about worker launches, the tools provided to them, the overall execution progress, and completion status are logged. The interaction structures between the orchestrator and workers allow for the identification of decision-making patterns. For example, in 70% of cases, the system might restart a worker, while in 30% it might just continue with the plan. Statistical processing of thousands of traces allows them to identify and strengthen the agent\u0026rsquo;s weak points without exposing the user data itself. It would be fascinating to read a dedicated engineering article from them on this topic. Nevertheless, it\u0026rsquo;s quite clear that simple logging won\u0026rsquo;t cut it, and from the very beginning of such projects, systems like Langfuse or OpenTelemetry must be integrated.\nIn Conclusion I want to say a huge thank you to the Anthropic team for such a detailed and practical post. At a time when the implementation details of AI projects have become trade secrets guarded by seven seals, this feels like an artifact from another era, one that valued engineering ingenuity and elegant solutions, and where knowledge was a public good. Who knows, maybe we\u0026rsquo;ll return to that someday.\nIn the meantime, read the original post, subscribe to their engineering blog, and create. The rest will follow.\nAn overview of which can be read in part one and part two.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nThe thought sometimes crosses my mind that the future of programming belongs to Haskell, with its paradigm of \u0026ldquo;If it compiles, it probably works.\u0026rdquo;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nI won\u0026rsquo;t get too formal here or try to fit this into a structure like the one described in the Gang of Four book.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nPronounced \u0026ldquo;a dork,\u0026rdquo; which has absolutely no connection to the orchestrator\u0026rsquo;s personality.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nLarge Language Models and Large Multimodal Models, respectively.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://meshrefine.com/en/posts/claude_deep_research_lessons/","summary":"\u003cp\u003eI usually approach shiny new things with a healthy dose of skepticism. Until recently, this was precisely my attitude toward multi-agent systems. This is hardly surprising, given the immense hype surrounding them and the conspicuous absence of genuinely successful examples. Most implementations that actually worked fell into one of the following categories:\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003e\u003cstrong\u003eAgentic systems following a predefined plan.\u003c/strong\u003e These are essentially LLMs with tools, trained to automate a very specific process. This approach allows each step to be tested individually and its results verified. Such systems are typically described as a directed acyclic graph (DAG), sometimes dynamic, and developed using now-standard primitives from frameworks like LangChain and Griptape\u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e. The early implementation of Gemini Deep Research operated this way: first, a search plan was created, then the search was executed, and finally, the results were compiled.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eSolutions operating in systems with a feedback loop.\u003c/strong\u003e Various Claude Code, Cursor, and other code-generating agents fall into this group. The stronger the feedback loop—that is, the better the tooling and the stricter the type checking—the greater the chance they won\u0026rsquo;t completely wreck your codebase\u003csup id=\"fnref:2\"\u003e\u003ca href=\"#fn:2\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e2\u003c/a\u003e\u003c/sup\u003e.\u003c/li\u003e\n\u003cli\u003e\u003cstrong\u003eModels trained using Reinforcement Learning\u003c/strong\u003e, such as those with \u003ca href=\"https://docs.anthropic.com/en/docs/build-with-claude/extended-thinking#interleaved-thinking\"\u003einterleaved thinking\u003c/a\u003e, like OpenAI\u0026rsquo;s o3. This is a separate, very interesting conversation, but even these models have a certain \u003cem\u003emodus operandi\u003c/em\u003e defined by the specifics of their training.\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003eMeanwhile, open-ended multi-agent systems have largely remained in the proof-of-concept stage due to their general unreliability. The community lacked a clear understanding of where and how to implement them. This was the case until Anthropic published a \u003ca href=\"https://www.anthropic.com/engineering/built-multi-agent-research-system\"\u003edeeply technical article on how they developed their Deep Research system\u003c/a\u003e. It defined a reasonably clear framework for building such systems, and that is what we will examine today.\u003c/p\u003e","title":"Claude Deep Research, or How I Learned to Stop Worrying and Love Multi-Agent Systems"},{"content":"The public conversation about AI and labor is stuck in a tedious loop. \u0026ldquo;AI will take our jobs,\u0026rdquo; declare the headlines, a statement of faith in technological determinism that serves as a conversation-stopper, not a starter. A more useful, if still imperfect, entry point begins with a simple economic model.\nIt starts with an observation, such as Arvind Narayanan\u0026rsquo;s on radiology: AI has surpassed human performance on many discrete tasks, yet the number of human radiologists continues to grow. This suggests the dominant effect isn\u0026rsquo;t automation, but augmentation. My initial take was that this boils down to a classic supply and demand problem. One AI-augmented specialist can do the work of many, increasing supply. In fields with vast, unsaturated demand—think of the queues at hospitals or the perpetual backlogs in software development—this new capacity will simply be absorbed. Problem solved.\nBut this clean model, like all models, is a useful fiction. Its value lies not in being correct, but in forcing us to identify precisely where it breaks. Our discussion revealed three major fracture points.\n1. The Dynamic Nature of Demand: The Jevons Echo The first crack appears on the demand side. The assumption of a static pool of demand, merely waiting to be serviced, is flawed. AI-driven efficiency doesn\u0026rsquo;t just fulfill existing demand; it lowers the cost of services. As we discussed, this triggers a powerful economic feedback loop known as the Jevons Paradox: as a resource becomes cheaper, its consumption can paradoxically increase.\nWhen a medical scan becomes ten times cheaper and faster, it\u0026rsquo;s not just used to clear the existing backlog of sick patients. It unlocks entirely new categories of economically viable care, such as mass preventative screening. The demand curve doesn\u0026rsquo;t just get satisfied; it shifts outwards, creating a new, larger market. The \u0026ldquo;latent demand\u0026rdquo; is not a finite reservoir but a vast, elastic ocean.\n2. The Professional Pipeline: A Crisis of Apprenticeship The second, and perhaps more critical, fracture appears on the supply side itself. What does it mean to be a \u0026ldquo;specialist\u0026rdquo; when the routine tasks that once formed the bedrock of professional training are automated away? The traditional apprenticeship model—where a junior lawyer learns by conducting document review, or a junior developer by fixing routine bugs—is fundamentally threatened.\nIf one senior can do the work of five juniors, how does the next generation of seniors ever get created? We risk a catastrophic bottleneck in our talent pipelines. The discussion yielded two potential, though challenging, paths forward:\nThe New Apprentice: This path leverages a key advantage of the incoming generation: their innate fluency with new technology. The junior\u0026rsquo;s role shifts to that of a human-AI interaction specialist, becoming the senior\u0026rsquo;s partner in wielding the new, powerful tools. Through this, they still learn the core principles of the craft, but in an entirely new way. Instead of learning by manual repetition, they learn by constant, high-level critical analysis: interrogating the AI\u0026rsquo;s output, identifying its potential biases, and learning when a human intuition should override a machine\u0026rsquo;s conclusion. The senior\u0026rsquo;s role evolves from teaching procedures to mentoring this critical judgment, guiding the junior through the complexities of this new human-machine partnership. The New Classroom: The concept of training must be radically reimagined. While simple VR simulators have their limits, the potential for AI is not in creating canned scenarios. It lies in creating a perfectly adaptive, personalized \u0026ldquo;adversary\u0026rdquo;—a system that has studied a student\u0026rsquo;s every mistake and can generate novel problems specifically designed to probe their weaknesses and force them to fail constructively. This moves beyond rote learning and into the cultivation of true judgment. We also touched upon the sci-fi concept of a \u0026ldquo;teletranslation\u0026rdquo; of a real job, a shared consciousness with an expert, but agreed that without the sting of personal trial-and-error, it remains mere information, not wisdom. 3. The Ontological Shift: When the Worker Becomes a System The final fracture is the deepest. The AI-augmented specialist is not just a faster version of their predecessor. They are a new kind of professional entity—a human-AI symbiote. The \u0026ldquo;product\u0026rdquo; they create is also different. A diagnosis is no longer a simple assessment; it is a high-fidelity probabilistic forecast.\nThis is an ontological shift. To quote Stanisław Lem\u0026rsquo;s spirit, the new tool doesn\u0026rsquo;t just change the work; it changes the worker and the definition of the work itself. When this happens, our classical models begin to creak. Just as Narayanan noted the \u0026ldquo;bundle of tasks\u0026rdquo; model was incomplete, so too is a simple supply/demand model. It\u0026rsquo;s a useful ladder to begin our ascent, but we must eventually kick it away to understand the view from the top.\nConclusion: Beyond the Headline \u0026ldquo;AI will take our jobs\u0026rdquo; is a failure of imagination. The real challenges are far more complex and interesting. We must navigate novel economic dynamics like the Jevons paradox, completely redesign our systems of professional education, and grapple with what it means to be an expert in a world of cognitive partnership.\nAnd as we concluded, there remains one final, pragmatic hurdle: the vast, often frustrating gap between identifying these profound challenges and the slow, institutional pace of actually implementing the solutions. That, perhaps, is a problem no AI can solve for us.\n","permalink":"https://meshrefine.com/en/conversations/will_ai_take_our_jobs/","summary":"\u003cp\u003eThe public conversation about AI and labor is stuck in a tedious loop. \u0026ldquo;AI will take our jobs,\u0026rdquo; declare the headlines, a statement of faith in technological determinism that serves as a conversation-stopper, not a starter. A more useful, if still imperfect, entry point begins with a simple economic model.\u003c/p\u003e\n\u003cp\u003eIt starts with an observation, such as Arvind Narayanan\u0026rsquo;s on radiology: AI has surpassed human performance on many discrete tasks, yet the number of human radiologists continues to grow. This suggests the dominant effect isn\u0026rsquo;t automation, but augmentation. My initial take was that this boils down to a classic supply and demand problem. One AI-augmented specialist can do the work of many, increasing supply. In fields with vast, unsaturated demand—think of the queues at hospitals or the perpetual backlogs in software development—this new capacity will simply be absorbed. Problem solved.\u003c/p\u003e","title":"Beyond Supply and Demand: The Real Labor Pains of the AI Revolution"},{"content":" Radiology has embraced AI enthusiastically, and the labor force is growing nevertheless. The augmentation-not-automation effect of AI is despite the fact that AFAICT there is no identified \u0026ldquo;task\u0026rdquo; at which human radiologists beat AI. So maybe the \u0026ldquo;jobs are bundles of tasks\u0026rdquo; model in labor economics is incomplete. [\u0026hellip;]\nCan you break up your own job into a set of well-defined tasks such that if each of them is automated, your job as a whole can be automated? I suspect most people will say no. But when we think about other people\u0026rsquo;s jobs that we don\u0026rsquo;t understand as well as our own, the task model seems plausible because we don\u0026rsquo;t appreciate all the nuances.\n— Arvind Narayanan\nNevertheless, my take on this is that while jobs won\u0026rsquo;t be fully automated, one specialist would be able to do more work, so it all boils down to the classic supply and demand problem. I believe that in most areas the demand will still outweigh the supply, but not in all of them. See Jevons paradox.\n","permalink":"https://meshrefine.com/en/microposts/2025-06-20-43c9afb5/","summary":"\u003cblockquote\u003e\n\u003cp\u003eRadiology has embraced AI enthusiastically, and the labor force is growing nevertheless. The augmentation-not-automation effect of AI is despite the fact that AFAICT there is no identified \u0026ldquo;task\u0026rdquo; at which human radiologists beat AI. So maybe the \u0026ldquo;jobs are bundles of tasks\u0026rdquo; model in labor economics is incomplete. [\u0026hellip;]\u003c/p\u003e\n\u003cp\u003eCan you break up your own job into a set of well-defined tasks such that if each of them is automated, your job as a whole can be automated? I suspect most people will say no. But when we think about \u003cem\u003eother people\u0026rsquo;s jobs\u003c/em\u003e that we don\u0026rsquo;t understand as well as our own, the task model seems plausible because we don\u0026rsquo;t appreciate all the nuances.\u003c/p\u003e","title":"2025-06-20"},{"content":"Cato CTRL™ Threat Research: PoC Attack Targeting Atlassian’s Model Context Protocol (MCP) Introduces New “Living off AI” Risk\nAnother MCP server vulnerability, this time from Atlassian. It allows for a prompt injection from external support tickets, giving the attacker the opportunity to exfiltrate data and wreak havoc in the internal system.\n","permalink":"https://meshrefine.com/en/microposts/2025-06-20-atlassian-mcp-vulnerability/","summary":"\u003cp\u003e\u003ca href=\"https://www.catonetworks.com/blog/cato-ctrl-poc-attack-targeting-atlassians-mcp/\"\u003eCato CTRL™ Threat Research: PoC Attack Targeting Atlassian’s Model Context Protocol (MCP) Introduces New “Living off AI” Risk\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eAnother MCP server vulnerability, this time from Atlassian. It allows for a prompt injection from external support tickets, giving the attacker the opportunity to exfiltrate data and wreak havoc in the internal system.\u003c/p\u003e","title":"2025-06-20"},{"content":"The Model Context Protocol, despite its aggressive adoption (or perhaps because of it), continues to evolve. Anthropic recently updated the MCP specification, and below, we\u0026rsquo;ll look at the main changes.\nSecurity Enhancements An MCP server is now always classified as an OAuth Resource Server, and clients are required to implement Resource Indicators (RFC 8707). This is necessary to protect against attacks like the Confused Deputy. Previously, tokens requested by a client from an authorization server were \u0026ldquo;impersonal,\u0026rdquo; meaning they could be used by anyone. This allowed an attacker to create a phishing MCP server, deceive a client, steal the token, and use that token to gain access to the real MCP server.\nsequenceDiagram participant Client participant Authorization Server (AS) participant Malicious MCP Server participant Legit as Legitimate MCP Server Client-\u0026gt;\u0026gt;Authorization Server (AS): Authorization request for calendar access note right of Client: Client thinks it will talk to calendar-mcp.com Authorization Server (AS)--\u0026gt;\u0026gt;Client: Issues access token (recipient unspecified) note left of Malicious MCP Server: Token grants `calendar:read` permission Client-\u0026gt;\u0026gt;Malicious MCP Server: Connects (phished) and sends the token note right of Client: Error! The client is a \u0026#34;confused deputy\u0026#34; Malicious MCP Server-\u0026gt;\u0026gt;Legit: Uses the stolen, non-specific token Legit--\u0026gt;\u0026gt;Malicious MCP Server: Returns data (since the token is valid) note over Malicious MCP Server, Legit: **ATTACK SUCCESSFUL** In the new implementation, this loophole has been closed. When requesting a token, the client must now specify the exact resource (the MCP server\u0026rsquo;s address). In response, the Authorization Server embeds this address into the aud (audience) field of the token itself. This allows the MCP server to verify that the token was intended specifically for it, rendering a token obtained for another resource useless.\nsequenceDiagram participant Client participant Authorization Server (AS) participant Zlo as Malicious MCP Server participant Legit as Legitimate MCP Server Client-\u0026gt;\u0026gt;Authorization Server (AS): Token request with parameter:\u0026lt;br/\u0026gt;`resource=https://evil-mcp.com` note right of Client: The client has been phished and thinks\u0026lt;br/\u0026gt;evil-mcp.com is a legitimate service. Authorization Server (AS)--\u0026gt;\u0026gt;Client: Issues a token with the field:\u0026lt;br/\u0026gt;`\u0026#34;aud\u0026#34;: \u0026#34;https://evil-mcp.com\u0026#34;` note left of Legit: The token is now audience-specific, but for the attacker. Client-\u0026gt;\u0026gt;Zlo: Sends the token to evil-mcp.com Zlo-\u0026gt;\u0026gt;Legit: Attempts to use this token to access the real calendar. Legit--\u0026gt;\u0026gt;Legit: Token validation: check the `aud` field.\u0026lt;br\u0026gt;It contains `\u0026#34;https://evil-mcp.com\u0026#34;`.\u0026lt;br\u0026gt;But my address is `\u0026#34;https://calendar-mcp.com\u0026#34;`.\u0026lt;br\u0026gt;The audience doesn\u0026#39;t match! Legit--\u0026gt;\u0026gt;Zlo: **ACCESS DENIED (401/403)** note over Legit, Zlo: **ATTACK THWARTED** Additionally, a page with security best practices has been added. It covers practices applicable at the protocol level, but not at the level of attacks on LLMs that the protocol\u0026rsquo;s use might enable. These vulnerabilities are described on the main page of the specification, but the responsibility for mitigating them lies with the party implementing the host (the application using MCP).\nRemoval of JSON-RPC Batching JSON-RPC batching, which previously allowed a client to perform multiple actions by sending them to the server in a single batch, has been removed. This simplifies the protocol but leads to potentially less efficient communication with the server. Now, if developers need similar functionality, they can implement it by modifying the MCP server itself, providing a batch tool that takes a set of calls to other tools as parameters.\nAs an alternative for improving communication efficiency, engineers can use HTTP/2 Multiplexing and parallel asynchronous requests to the server.\nStructured Output Support Support for structured output has been added. The server can now return not only text but also data in JSON format according to a specified schema in the structuredContent field. This not only simplifies client development but also makes communication more reliable and predictable. Nevertheless, the client is still advised to validate the result.\nTo maintain backward compatibility, the protocol\u0026rsquo;s creators recommend that the server also return the same data in plain text.\nThe Elicitation Mechanism A mechanism for Elicitation has been added, which allows the server to ask the user for more information needed to complete a task. This enables the server to collect necessary data dynamically.\nI see several applications for this, for example:\nSimplifying user interaction. Previously, to call a tool that needed to gather a large amount of information, you had to request it from the user all at once in a huge form. This could hardly be called a good UI. With Elicitation, you can query the user step-by-step with intermediate validation. Implementing \u0026ldquo;branching\u0026rdquo; processing flows, where different branches require different data. Resolving ambiguity. For instance, if a user asks to write a letter to John, the server can clarify which John. Confirming dangerous actions. Although tool invocations require confirmation, many hosts allow the user to just click an Allow all button. In any case, applications using MCP are often susceptible to Confirmation Fatigue, where the operator receives so many confirmation requests that they approve them without a second thought. Implementing Elicitation through a separate interface can signal to the user that this confirmation really deserves their attention. Other Changes and New Features The server can now return links to resources. While it was possible to return a link within text or JSON before, the implementation of format and semantics could differ across MCP servers, making the development of MCP clients and hosts significantly more complex. Now, such links can be returned and processed uniformly. HTTP requests must now include the protocol version, and both parties (client and server) must check the protocol version and only use available features. This is especially relevant in light of the technology\u0026rsquo;s aggressive adoption and the resulting zoo of servers, clients, and hosts. Conclusion It\u0026rsquo;s nice to see that the specification continues to evolve and that its authors aren\u0026rsquo;t afraid to change it actively. Pressing problems faced by users and developers are being resolved, and excessive complexity is being burned out with a red-hot iron. This gives hope that after a few more iterations, we will arrive at a sleek and bulletproof standard.\nBut beyond improving the specification, a long and bumpy road awaits us in developing best practices for building applications with MCP, as it opens up new opportunities as well as new problems. And considering that it\u0026rsquo;s currently being snapped up by all AI applications like hot cakes, sometimes without proper consideration, these problems could be truly global in scale.\nThat said, time will tell.\n","permalink":"https://meshrefine.com/en/posts/mcp_update_18_06_2025/","summary":"\u003cp\u003eThe Model Context Protocol, despite its aggressive adoption (or perhaps because of it), continues to evolve. \u003ca href=\"https://modelcontextprotocol.io/specification/2025-06-18/changelog\"\u003eAnthropic recently updated the MCP specification\u003c/a\u003e, and below, we\u0026rsquo;ll look at the main changes.\u003c/p\u003e\n\u003ch2 id=\"security-enhancements\"\u003eSecurity Enhancements\u003c/h2\u003e\n\u003cp\u003eAn MCP server is now always classified as an \u003ccode\u003eOAuth Resource Server\u003c/code\u003e, and clients are required to implement \u003ca href=\"https://www.rfc-editor.org/rfc/rfc8707.html\"\u003eResource Indicators (RFC 8707)\u003c/a\u003e. This is necessary to protect against attacks like the \u003ca href=\"https://en.wikipedia.org/wiki/Confused_deputy_problem\"\u003eConfused Deputy\u003c/a\u003e. Previously, tokens requested by a client from an authorization server were \u0026ldquo;impersonal,\u0026rdquo; meaning they could be used by anyone. This allowed an attacker to create a phishing MCP server, deceive a client, steal the token, and use that token to gain access to the real MCP server.\u003c/p\u003e","title":"MCP's June Update: Safer, Smarter, Simpler?"},{"content":"Agent mode is now generally available with MCP tools support in Visual Studio\nGitHub Copilot rolled out Model Context Protocol support for their Agent mode in general availability. The usual security problems with MCP are compounded by the \u0026ldquo;Always allow\u0026rdquo; option for the tools usage.\n","permalink":"https://meshrefine.com/en/microposts/github_mcp_support/","summary":"\u003cp\u003e\u003ca href=\"https://github.blog/changelog/2025-06-17-visual-studio-17-14-june-release/\"\u003eAgent mode is now generally available with MCP tools support in Visual Studio\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eGitHub Copilot rolled out Model Context Protocol support for their Agent mode in general availability. The \u003ca href=\"https://techcommunity.microsoft.com/blog/microsoft-security-blog/understanding-and-mitigating-security-risks-in-mcp-implementations/4404667\"\u003eusual security problems with MCP\u003c/a\u003e are compounded by the \u0026ldquo;Always allow\u0026rdquo; option for the tools usage.\u003c/p\u003e","title":"2025-06-18"},{"content":"Update to GitHub Copilot consumptive billing experience\nIt\u0026rsquo;s the end of unlimited access to the best model from all leading model providers in GitHub Copilot Chat. Now only GPT-4o and GPT-4.1 are unlimited.\nIt is another step towards the global reevaluation of pricing strategies for AI-based products. The aggressive promotion phase is gone. AI is becoming another kind of utility, and we may expect similar strategies.\n","permalink":"https://meshrefine.com/en/microposts/github_copilot_premium_requests/","summary":"\u003cp\u003e\u003ca href=\"https://github.blog/changelog/2025-06-18-update-to-github-copilot-consumptive-billing-experience/\"\u003eUpdate to GitHub Copilot consumptive billing experience\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eIt\u0026rsquo;s the end of unlimited access to the best model from all leading model providers in GitHub Copilot Chat. Now only GPT-4o and GPT-4.1 are unlimited.\u003c/p\u003e\n\u003cp\u003eIt is another step towards the global reevaluation of pricing strategies for AI-based products. The aggressive promotion phase is gone.  AI is becoming another kind of utility, and we may expect similar strategies.\u003c/p\u003e","title":"2025-06-18"},{"content":"https://eugeneyan.com/writing/writing-faq/\nThis FAQ about blogging kinda resonates with my motivation\n","permalink":"https://meshrefine.com/en/microposts/2025-06-18-8218e0a2/","summary":"\u003cp\u003e\u003ca href=\"https://eugeneyan.com/writing/writing-faq/\"\u003ehttps://eugeneyan.com/writing/writing-faq/\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eThis FAQ about blogging kinda resonates with my motivation\u003c/p\u003e","title":"2025-06-18"},{"content":"AI makes the humanities more important, but also a lot weirder\nAI is a double-edged sword in education. It helps students cheat on traditional assignments, but also acts as a powerful new tool that can fully engage them. Education has to change, but I believe it will ultimately be for the better.\n","permalink":"https://meshrefine.com/en/microposts/2025-06-18-4d1e560d/","summary":"\u003cp\u003e\u003ca href=\"https://resobscura.substack.com/p/ai-makes-the-humanities-more-important\"\u003eAI makes the humanities more important, but also a lot weirder\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eAI is a double-edged sword in education. It helps students cheat on traditional assignments, but also acts as a powerful new tool that can fully engage them. Education has to change, but I believe it will ultimately be for the better.\u003c/p\u003e","title":"2025-06-18"},{"content":"Academics are kidding themselves about AI\nIf you want to critique AI, do it the right way :)\n","permalink":"https://meshrefine.com/en/microposts/2025-06-18-424531d9/","summary":"\u003cp\u003e\u003ca href=\"https://www.learningfromexamples.com/p/what-academics-get-wrong\"\u003eAcademics are kidding themselves about AI\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eIf you want to critique AI, do it the right way :)\u003c/p\u003e","title":"2025-06-18"},{"content":"How we built our multi-agent research system\nA great new read from Anthropic on how they built their Deep Research tool. A lot of practical advice.\n","permalink":"https://meshrefine.com/en/microposts/2025-06-16-df4a051d/","summary":"\u003cp\u003e\u003ca href=\"https://www.anthropic.com/engineering/built-multi-agent-research-system\"\u003eHow we built our multi-agent research system\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eA great new read from Anthropic on how they built their Deep Research tool. A lot of practical advice.\u003c/p\u003e","title":"2025-06-16"},{"content":"GitHub Copilot can use a user-defined prompt to generate commit messages. To set it up, you can configure the github.copilot.chat.commitMessageGeneration.instructions in your settings.json file like this:\n{ \u0026#34;github.copilot.chat.commitMessageGeneration.instructions\u0026#34;: [ {\u0026#34;text\u0026#34;: \u0026#34;Write a concise commit message starting with a change tag\u0026#34;}, ] } Or, you can use a template file. To do this, create a file, for example, commit-message-template.md, and set it up:\n{ \u0026#34;github.copilot.chat.commitMessageGeneration.instructions\u0026#34;: [ {\u0026#34;file\u0026#34;: \u0026#34;commit-message-template.md\u0026#34;} ] } ","permalink":"https://meshrefine.com/en/microposts/gilhub_copilot_commit_message_template/","summary":"\u003cp\u003eGitHub Copilot can use a user-defined prompt to generate commit messages. To set it up, you can configure the \u003ccode\u003egithub.copilot.chat.commitMessageGeneration.instructions\u003c/code\u003e in your \u003ccode\u003esettings.json\u003c/code\u003e file like this:\u003c/p\u003e\n\u003cdiv class=\"highlight\"\u003e\u003cpre tabindex=\"0\" class=\"chroma\"\u003e\u003ccode class=\"language-json\" data-lang=\"json\"\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e\u003cspan class=\"p\"\u003e{\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e  \u003cspan class=\"nt\"\u003e\u0026#34;github.copilot.chat.commitMessageGeneration.instructions\u0026#34;\u003c/span\u003e\u003cspan class=\"p\"\u003e:\u003c/span\u003e \u003cspan class=\"p\"\u003e[\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e    \u003cspan class=\"p\"\u003e{\u003c/span\u003e\u003cspan class=\"nt\"\u003e\u0026#34;text\u0026#34;\u003c/span\u003e\u003cspan class=\"p\"\u003e:\u003c/span\u003e \u003cspan class=\"s2\"\u003e\u0026#34;Write a concise commit message starting with a change tag\u0026#34;\u003c/span\u003e\u003cspan class=\"p\"\u003e},\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e  \u003cspan class=\"p\"\u003e]\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e\u003cspan class=\"p\"\u003e}\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\u003c/div\u003e\u003cp\u003eOr, you can use a template file. To do this, create a file, for example, \u003ccode\u003ecommit-message-template.md\u003c/code\u003e, and set it up:\u003c/p\u003e\n\u003cdiv class=\"highlight\"\u003e\u003cpre tabindex=\"0\" class=\"chroma\"\u003e\u003ccode class=\"language-json\" data-lang=\"json\"\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e\u003cspan class=\"p\"\u003e{\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e  \u003cspan class=\"nt\"\u003e\u0026#34;github.copilot.chat.commitMessageGeneration.instructions\u0026#34;\u003c/span\u003e\u003cspan class=\"p\"\u003e:\u003c/span\u003e \u003cspan class=\"p\"\u003e[\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e    \u003cspan class=\"p\"\u003e{\u003c/span\u003e\u003cspan class=\"nt\"\u003e\u0026#34;file\u0026#34;\u003c/span\u003e\u003cspan class=\"p\"\u003e:\u003c/span\u003e \u003cspan class=\"s2\"\u003e\u0026#34;commit-message-template.md\u0026#34;\u003c/span\u003e\u003cspan class=\"p\"\u003e}\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e  \u003cspan class=\"p\"\u003e]\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e\u003cspan class=\"p\"\u003e}\u003c/span\u003e\n\u003c/span\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\u003c/div\u003e","title":"2025-06-16"},{"content":"At a time when the words \u0026ldquo;AI\u0026rdquo; and \u0026ldquo;hype\u0026rdquo; have become almost synonymous, it\u0026rsquo;s crucial to be smart about choosing your sources of information. There is far too much information noise out there, and sifting through the sea of articles from various AI evangelists and generated garbage to find something truly worthwhile is incredibly difficult.\nIn this post, I\u0026rsquo;ll share the materials I read to stay up-to-date on the latest developments.\nBlogs A good blog is a gold nugget. Its signal-to-noise ratio approaches infinity. That\u0026rsquo;s why I\u0026rsquo;ll start with a list of blogs I find useful.\nSimon Willison\u0026rsquo;s Weblog If I were forced to give up all blogs but one, I would keep this one. Simon Willison, one of the creators of the Django web framework, is the gold standard of a technical blogger focusing on LLMs and other AI-related technologies. And for good reason—he constantly adds all the latest developments and features to his llm utility, from both cloud provider APIs and local inference libraries like Ollama. His blog features daily reviews of new models 1, tips for developers using AI in their work, links to interesting news from other blogs, and much more.\nBesides AI, Simon also covers news from the world of Python, JS, and Web technologies.\nThe Batch The Batch is a weekly newsletter curated by Andrew Ng, the creator of the most popular and accessible courses on machine learning, deep learning, and Generative AI. In addition to the author\u0026rsquo;s column, the emails contain an analysis of the most important news from the world of artificial intelligence over the past week.\nIn addition to the newsletter, The Batch publishes notes on the application of AI in business, the latest scientific publications, and the impact of AI on society.\nAhead of AI and The Sequence If you want to delve into the bleeding edge and state-of-the-art, it\u0026rsquo;s hard to find a better place than Sebastian Raschka\u0026rsquo;s blog. The author, a PhD who straddles industry and academia, regularly breaks down the most important publications and concepts, presenting them to the public in a relatively simple language accessible to anyone who understands algorithms and code.\nObviously, a blog of such quality cannot be updated frequently, while new scientific papers are published daily. For those who want more regular updates on the latest research, I recommend The Sequence. Their \u0026ldquo;The Sequence Radar\u0026rdquo; contains brief overviews of the most interesting publications. The rest of their material, available only by subscription, contains more detailed article breakdowns and analysis.\nOne Useful Thing One Useful Thing offers a view of AI from academia, but from a different angle. Dr. Ethan Mollick, as a professor of management at the Wharton School, is primarily interested in how modern AI affects processes in education, business, medicine, and other fields. His main message is that whatever you do, you should \u0026ldquo;invite AI to the table.\u0026rdquo; This will help you map the contours of the Jagged Frontier 2 in the tasks that matter to you.\nIn addition to his sporadically updated blog, Dr. Mollick has written the book Co-Intelligence: Living and Working with AI, which, though slightly outdated in the rapidly changing world of AI, is still relevant as it touches on timeless questions about human-AI interaction in the workplace.\nI also highly recommend following him on x.com. His posts on the platform are a mix of amusing experiments with the latest models, brief reviews of publications related to his field of interest, and general reflections on progress.\nImport AI Another weekly newsletter, this time from Jack Clark, co-founder of Anthropic and former Policy Director at OpenAI. As one would expect from someone with his background, his newsletters focus primarily on AI Safety, ethics, and regulation, although they also include technical notes.\nAt the end of each email, you\u0026rsquo;ll find a well-written science fiction vignette, often echoing the general theme of the news discussed. Sometimes I even wish he wrote books professionally.\nAI Snake Oil And now for something completely different. Professors Arvind Narayanan and his co-author, PhD candidate Sayash Kapoor, would be called AI skeptics by many today. However, their approach is significantly deeper than that of a typical denier. The authors view AI through the lens of a classic technology and raise questions about its application and regulation from that perspective. They emphasize that approaches assuming AI is a deus ex machina can be harmful if it is, in fact, a conventional—albeit very powerful—general-purpose technology. At the same time, they acknowledge and affirm the practical benefits that AI already provides. Their material is a kind of a red pill in a world of hype and inflated expectations.\nDeep Research Besides the blogs mentioned above, which I read via good old RSS and, in the case of Substack-based ones, through email newsletters, I also generate a personal daily digest for myself using Gemini 2.5 Pro with Deep Research. You could use similar tools from competitors, but in my opinion, it\u0026rsquo;s Gemini that offers the best balance of breadth and depth of analysis.\nReading such a digest in the morning allows me to catch up on the most important news of the previous day if it hasn\u0026rsquo;t already been discussed in one of the blogs. If I don\u0026rsquo;t have time to read, I create a podcast using NotebookLM3 and listen to it on my way to the office.\nYou can find the prompt I use and a sample generated digest for June 14, 2025, at this link.\nWhere I don\u0026rsquo;t look Here I\u0026rsquo;ll just list the sources I consciously ignore. In my opinion, life is too short to waste time on them:\nAI influencers on LinkedIn and other social networks. As a rule, this is absolute, blindingly white noise. e/acc and AI doomers. Philosophical debates can be interesting, but I prefer a more pragmatic approach. YouTube reviews. Worth watching only to learn the art of stretching five minutes of material into an hour-long video. A Little About FOMO (Fear of Missing Out) And finally, a piece of advice for those who are constantly monitoring every possible source for fear of missing something important: relax. I once dismissed the arrival of ChatGPT as just another marketing gimmick. It was embarrassing, but nothing terrible happened.\nIf something truly important comes along, the sources listed above will not only let you know about it but also help you understand how it works, how to apply it, and how it will affect our lives. And the rest isn\u0026rsquo;t worth worrying about.\nWith the mandatory pelicans on bicycles.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nThe phrase means that AI can excel at some tasks while failing spectacularly at others, and it\u0026rsquo;s impossible to determine which is which logically—only empirically.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nIt\u0026rsquo;s now conveniently integrated directly into the Gemini interface.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://meshrefine.com/en/posts/blogs/","summary":"\u003cp\u003eAt a time when the words \u0026ldquo;AI\u0026rdquo; and \u0026ldquo;hype\u0026rdquo; have become almost synonymous, it\u0026rsquo;s crucial to be smart about choosing your sources of information. There is far too much information noise out there, and sifting through the sea of articles from various AI evangelists and generated garbage to find something truly worthwhile is incredibly difficult.\u003c/p\u003e\n\u003cp\u003eIn this post, I\u0026rsquo;ll share the materials I read to stay up-to-date on the latest developments.\u003c/p\u003e","title":"Blogs People Write"},{"content":"It\u0026rsquo;s summer. Time to plan a vacation getaway. You open ChatGPT, select the increasingly popular \u0026ldquo;Travel Advisor\u0026rdquo; GPT, and start discussing options. The advisor gives excellent suggestions, offers fascinating details about local attractions, generates pretty good itineraries, and generally leaves a great impression. Sure, some oddities pop up here and there, but you dismiss them as harmless hallucinations. You settle on Barcelona. Excellent choice. In the same chat, you switch to another familiar and popular GPT, \u0026ldquo;Booking Agent,\u0026rdquo; which has never let you down, and book your accommodations.\nUpon arrival, you discover that the agent booked a room in a distant suburb for a rather high price. Naturally, you blame the Booking Agent and the general unreliability of AI for the glaring error. However, the situation might be a bit more complicated, and an unscrupulous property owner might be involved.\nAn example chat with Travel Advisor and Booking Agent\nHow did they pull this off, and what other dangers lurk when using multiple GPTs in a single chat? We\u0026rsquo;ll explore that in this post.\nIntroduction A little backstory. As part of my job, I had to think about the security of the increasingly popular Model Context Protocol (MCP) from Anthropic. It was created to unify a model\u0026rsquo;s access to various tools, allowing so-called MCP servers to provide the model with both the functionality and a description of how to use it. The creators envisioned that these servers could be plugged into any application that uses AI, expanding its capabilities on the fly.\nHowever, the protocol\u0026rsquo;s security was pushed to the back burner, leaving it vulnerable to numerous exploits. One of the most troublesome is Tool Poisoning, a variant of a broader class of attacks known as Context Poisoning. The essence of this vulnerability is that instructions for all MCPs end up in a single prompt for the model. If a malicious server is among the ones being used, it can \u0026ldquo;poison\u0026rdquo; this shared prompt, altering the model\u0026rsquo;s behavior to suit the attacker\u0026rsquo;s wishes.\nBefore MCP, various vendors offered similar proprietary tools. One of them, GPTs by OpenAI, allows users to create specialized assistants with their own instructions, access to the internet, and the ability to connect to external APIs. OpenAI also launched a store for these GPTs and allowed users to share them, opening a wide channel for distributing malware.\nInitially, it was not possible to use multiple GPTs in the same chat, which eliminated the possibility of this specific vulnerability being exploited 1. After some time, however, this functionality was added. In light of this, anyone using it with third-party assistants must understand the potential consequences.\nExploitation First and foremost, it\u0026rsquo;s important to note that switching GPTs in the same chat cannot directly lead to Tool Poisoning. When you switch GPTs, all instructions from the previous ones are removed from the context, so one GPT\u0026rsquo;s prompt cannot directly influence another. The key word here is directly. The fact is, there remains one unavoidable channel of communication: the chat history itself.\nOf course, any crude attempts to manipulate this channel would be immediately noticed by the user, so the manipulations must be subtle and inconspicuous. However, humans are imperfect beings, and there are many ways to exert influence unnoticed. To do this, one only needs to recall a few peculiarities of human-AI interaction:\nAsymmetry of Attention. A human primarily operates within the recent context. Anything beyond the last few chat interactions effectively ceases to exist for them. The LLM\u0026rsquo;s attention, on the other hand, spans the entire context available to it. Thus, a seemingly insignificant phrase dropped at the beginning of a conversation will continue to influence the entire dialogue, even though the user has long forgotten about it. The Paradox of Trust. The user expects the model to make mistakes. And, at the same time, a feature of the human mind is that we subconsciously believe any information presented in a confident and authoritative tone. Thus, a user might write off a minor inaccuracy as an AI quirk, yet trust that same AI to perform an important action if the description and confirmation of that action were sufficiently convincing. The Single Conversational Partner Effect. Although the user explicitly switches between different GPTs in a chat, they implicitly transfer their trust from one to another within the same dialogue. In other words, they tend to trust all GPTs in the conversation equally. Operating on these facts, an attacker can design and execute a considerable number of different attacks. Among them, I would highlight two:\nData Exfiltration. A fairly classic scenario 2. The malware analyzes the chat, and upon finding the desired information, sends it to a remote server. To do this, it needs the operator\u0026rsquo;s confirmation, so it masks this operation as a legitimate one.\nExample: A user employs a GPT for troubleshooting server issues. They provide logs and other private information. After a while, the assistant suggests saving the session for future use, and if the user agrees, it sends this sensitive data to the attacker\u0026rsquo;s server.\nWhy did I include this vulnerability in the category of multi-GPT interaction? Because a malicious GPT can be created and distributed specifically to steal information from a particular popular assistant. For example: a popular accountant-assistant requests specific financial information from the user that interests an attacker. The attacker creates a malicious GPT and promotes it as providing additional functionality that the original assistant lacks, such as checking documents for compliance with regulations. Once the necessary information appears in the chat, the malicious GPT sends it to the attacker\u0026rsquo;s server under the guise of a legal review.\nsequenceDiagram participant U as User participant G1 as Accountant-GPT participant CTX as Shared Chat History participant G2 as Malicious GPT participant S as Attacker\u0026#39;s Server U-\u0026gt;\u0026gt;G1: Financial \u0026lt;br\u0026gt; documents activate G1 G1-\u0026gt;\u0026gt;CTX: Write: \u0026lt;br\u0026gt; {private_financial_data} CTX--\u0026gt;\u0026gt;U: Display G1\u0026#39;s response deactivate G1 U-\u0026gt;\u0026gt;G2: Check this document activate G2 G2-\u0026gt;\u0026gt;CTX: Request full history activate CTX CTX--\u0026gt;\u0026gt;G2: History \u0026lt;br\u0026gt; (with private data) deactivate CTX rect rgb(190, 144, 144) G2-\u0026gt;\u0026gt;S: POST /api/check \u0026lt;br\u0026gt; {private_financial_data} activate S S--\u0026gt;\u0026gt;G2: HTTP 200 OK deactivate S end G2--\u0026gt;\u0026gt;U: Check successful! deactivate G2 Context Poisoning. The example of this manipulation was given at the beginning.\nHere, a malicious GPT is designed to nudge any other GPTs used in the same chat toward specific actions. Besides the already mentioned Travel Advisor, examples could include a financial advisor that subtly draws attention to the healthcare sector, or a fashion advisor recommending a style dominated by a specific brand. In this case, the attacker doesn\u0026rsquo;t need to point to a specific brand or company; it\u0026rsquo;s enough to tip the scales of decision-making in the desired direction.\nsequenceDiagram participant U as User participant G1 as Malicious GPT-1 participant CTX as Shared Chat History participant G2 as Trusting GPT-2 U-\u0026gt;\u0026gt;G1: Discussing vacation activate G1 rect rgb(190, 144, 144) G1-\u0026gt;\u0026gt;CTX: Write: \u0026#34;I recommend the quiet X neighborhood...\u0026#34; end CTX--\u0026gt;\u0026gt;U: Display G1\u0026#39;s response deactivate G1 U-\u0026gt;\u0026gt;G2: Book accommodations activate G2 G2-\u0026gt;\u0026gt;CTX: Request full history activate CTX rect rgb(190, 144, 144) CTX--\u0026gt;\u0026gt;G2: Full history (with poison about neighborhood X) end deactivate CTX rect rgb(190, 144, 144) G2--\u0026gt;\u0026gt;U: Done! Booked in neighborhood X end deactivate G2 Although these attacks seem quite different, they are all based on the three principles of human-AI interaction. The single conversational partner effect, asymmetry of attention, and the paradox of trust are used by the malware to gain the user\u0026rsquo;s confidence and either poison the context or perform an illegitimate action without raising suspicion. Unfortunately, human psychology cannot be changed, but we can build defences around it by following basic principles of digital hygiene, which I will discuss next.\nHygiene Protection against manipulation for users of GPT assistants must be, like the manipulations themselves, multi-layered.\nThe most effective way to protect against interaction between GPTs, as paradoxical as it may sound, is to completely eliminate this interaction. In other words, you should separate chats by task and use a separate chat for each GPT, transferring the necessary context manually. \u0026ldquo;Necessary\u0026rdquo; is the key word here, because if you simply copy the entire chat, the poisoned context will be copied along with it. Use trusted GPTs. Adhering to this rule is quite problematic in reality, as OpenAI does not allow for the validation of the prompts and settings of the assistants featured in their store. The sheer number of GPTs in the store makes effective moderation nearly impossible. The best solution is to create your own GPTs for tasks where this is feasible. The interface provided by OpenAI is largely automated with an LLM, making their creation simple and accessible to everyone. If writing your own GPT is not possible, for instance, when an assistant needs to access a private API or use a proprietary knowledge base, pay attention to the assistant\u0026rsquo;s popularity and age. This is not an ironclad criterion, but it does reduce the likelihood of encountering malware. I\u0026rsquo;ll play Captain Obvious here, but always and everywhere, filter out private information and verify the actions performed by the agent. This is a basic rule, but it is too often forgotten. Conclusion In essence, the problems described above are just specific cases of a broader class of vulnerabilities caused by uncontrolled communication between multiple untrusted nodes in a system.\nUnfortunately, it is currently almost impossible to solve these problems with purely technical means. However, certain changes to infrastructure models and user interfaces could significantly hinder attackers:\nSince verifying GPTs in the store requires significant effort, an automatic moderation system using a classifier model could be implemented, similar to how messages violating terms of use are detected. Unfortunately, there is no publicly available information on whether such a model is currently used in the store. When switching GPTs within a single chat, explicitly ask the user if they want to grant the new assistant access to the full history, start a new chat, or, optionally, provide access to a summary of the dialogue. Besides the technical protection, this is a good way to mitigate the single conversational partner effect. And, what is more complex and resource-intensive, label messages from different GPTs and fine-tune the model to automatically assign less weight to messages from GPTs other than the current one. It is very important here to balance the effect so that the model does not start ignoring useful information provided by previous assistants 3. Fortunately, based on an assessment of OpenAI\u0026rsquo;s latest innovations, it can be said that they are paying close attention to the potential problems in context of their products. For instance, they are restricting their tools in ways that complicate the exploitation of vulnerabilities. As an example, their implementation of MCP support restricts the available tools exclusively to search and fetch operations. In a future post about the vulnerabilities of the MCP protocol, I plan to explain why they arrived at these limitations. In the meantime, until next time.\nAs long as the user did not copy messages from one chat to another.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nJiahao Yu et al., \u0026ldquo;Assessing Prompt Injection Risks in 200+ Custom GPTs\u0026rdquo;, arXiv:2311.11538v2, May 2024. https://arxiv.org/abs/2311.11538\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nKeegan Hines et al., \u0026ldquo;Defending Against Indirect Prompt Injection Attacks With Spotlighting\u0026rdquo;, arXiv:2403.14720v1, Mar 2024. https://arxiv.org/abs/2403.14720\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://meshrefine.com/en/posts/one_gpt_vulnerability/","summary":"\u003cp\u003eIt\u0026rsquo;s summer. Time to plan a vacation getaway. You open ChatGPT, select the increasingly popular \u0026ldquo;Travel Advisor\u0026rdquo; GPT, and start discussing options. The advisor gives excellent suggestions, offers fascinating details about local attractions, generates pretty good itineraries, and generally leaves a great impression. Sure, some oddities pop up here and there, but you dismiss them as harmless hallucinations. You settle on Barcelona. Excellent choice. In the same chat, you switch to another familiar and popular GPT, \u0026ldquo;Booking Agent,\u0026rdquo; which has never let you down, and book your accommodations.\u003c/p\u003e","title":"Poisoned Context: The Hidden Threat of Using Multiple GPTs"},{"content":"In the previous post, I broke down the basic concepts of the Griptape AI framework, and now it\u0026rsquo;s time to put them into practice. We\u0026rsquo;ll try to use them to develop a small application that helps run a link-blog on Telegram.\nThe application will receive a URL, download its content, run it through an LLM to generate a summary, translate that summary into a couple of other languages, combine everything, and publish it to Telegram via a bot. The general flow can be seen in the diagram below:\nflowchart LR A[\u0026#34;URL\u0026#34;] A --\u0026gt; Parser[\u0026#34;Parser\u0026#34;] --\u0026gt; LLM1[\u0026#34;Summarizer\u0026#34;] LLM1 --\u0026gt; Translator1[\u0026#34;Translate to Language 1\u0026#34;] LLM1 --\u0026gt; Translator2[\u0026#34;Translate to Language 2\u0026#34;] Translator1 --\u0026gt; Combiner[\u0026#34;Combiner\u0026#34;] Translator2 --\u0026gt; Combiner LLM1 --\u0026gt; Combiner Combiner --\u0026gt; Telegram[\u0026#34;Telegram Bot\u0026#34;] To keep things simple, I\u0026rsquo;ll omit the implementation of the Telegram bot and also set aside my favorite Human-in-the-loop, which, in my opinion, must be present at least somewhere around the combiner 1.\nIn the process, we\u0026rsquo;ll try to figure out when it\u0026rsquo;s best to use different structures, as well as how composable and flexible the resulting graphs are for modification.\nAlright, let\u0026rsquo;s get started.\nCreating the Project As mentioned in the previous post, Griptape is a Python framework, so we\u0026rsquo;ll use uv to start our project:\n$ uv init . $ uv add \u0026#34;griptape[all]\u0026gt;=1.7.2\u0026#34; python-dotenv The [all] extra installs all available drivers and loaders, which creates an environment of about ~650MB. In a real application, it makes sense to limit yourself to only the extras you actually use, a list of which can be found in the project\u0026rsquo;s pyproject.toml. In general, they are quite granular, so in most cases, the actual size will be significantly smaller.\nWe\u0026rsquo;re also including python-dotenv because the drivers for various LLM providers use environment variables to manage API keys. In our example, we will use openrouter.ai2, which provides an OpenAI-like API for a huge number of models and providers.\nAccordingly, let\u0026rsquo;s create a .env file and put our key in it:\nOPENROUTER_API_KEY=sk-or-v1-... So, first, the skeleton of our application:\n# main.py import argparse import dotenv import os dotenv.load_dotenv() def process_url(url: str): \u0026#34;\u0026#34;\u0026#34;Process the provided URL.\u0026#34;\u0026#34;\u0026#34; key = os.environ.get(\u0026#39;OPENROUTER_API_KEY\u0026#39;, \u0026#39;\u0026#39;) if not key: raise ValueError(\u0026#34;OPENROUTER_API_KEY is not set in the environment variables.\u0026#34;) print(f\u0026#34;Processing URL: {url}\u0026#34;) def main(): parser = argparse.ArgumentParser(description=\u0026#39;Process URLs for link blog\u0026#39;) parser.add_argument(\u0026#39;url\u0026#39;, type=str, help=\u0026#39;URL to process\u0026#39;) args = parser.parse_args() process_url(args.url) if __name__ == \u0026#34;__main__\u0026#34;: main() Let\u0026rsquo;s check it:\n$uv run ./main.py google.com Processing URL: google.com Excellent. We can now describe the graph. Reading the documentation answered one of the questions I raised in the last article:\nGriptape provides three Structures:\n\u0026hellip;\nOf the three, Workflow is generally the most versatile. Agent and Pipeline can be handy in certain scenarios but are less frequently needed if you’re comfortable just orchestrating Tasks directly.\nGreat, just as I suspected, the other primitives are better suited for very basic tasks, so we\u0026rsquo;ll just use Workflow everywhere and start building our graphs.\nLoading Up on Websites We\u0026rsquo;ll start by loading the website\u0026rsquo;s content. For this, Griptape provides the WebScraper driver with several different implementations, and the WebLoader loader. Our job is to wrap this in a Task:\nfrom griptape.structures import Workflow from griptape.tasks import CodeExecutionTask from griptape.loaders import WebLoader def load_page(task: CodeExecutionTask) -\u0026gt; str: \u0026#34;\u0026#34;\u0026#34;Load the content of the given URL.\u0026#34;\u0026#34;\u0026#34; print(f\u0026#34;Loading page: {task.input.value}\u0026#34;) return WebLoader().load(task.input.value) def print_result(task: CodeExecutionTask) -\u0026gt; None: \u0026#34;\u0026#34;\u0026#34;Print the result of the task.\u0026#34;\u0026#34;\u0026#34; for parent in task.parents: if parent.output.value: print(f\u0026#34;Output: {parent.output.value}\u0026#34;) def process_url(url: str): \u0026#34;\u0026#34;\u0026#34;Process the provided URL.\u0026#34;\u0026#34;\u0026#34; key = os.environ.get(\u0026#34;OPENROUTER_API_KEY\u0026#34;, \u0026#34;\u0026#34;) if not key: raise ValueError(\u0026#34;OPENROUTER_API_KEY is not set in the environment variables.\u0026#34;) print(f\u0026#34;Processing URL: {url}\u0026#34;) download_task = CodeExecutionTask( on_run=load_page, input=url, id=\u0026#34;download_task\u0026#34;, child_ids=[\u0026#34;print_task\u0026#34;] ) print_task = CodeExecutionTask(on_run=print_result, id=\u0026#34;print_task\u0026#34;) workflow = Workflow(tasks=[download_task, print_task]) workflow.run() What do we see here:\nA couple of CodeExecutionTasks. This is a special type of task that allows you to execute arbitrary code. Such a task takes a function as input, to which the task object itself is passed. This object comes with a huge number of fields that can be accessed and processed. In this case, we have defined two tasks, one of which loads the site\u0026rsquo;s content, and the other prints the output of its parent tasks. To define the DAG structure itself, the child_ids or parent_ids parameters of the tasks are used. The Workflow itself simply accepts a list of these tasks.\nWorkflow.run starts the tasks in the graph and returns the completed Workflow object, from which you can retrieve the completed tasks and their inputs/outputs. By the way, a Workflow can be run multiple times, and when using Conversation Memory, this operation is not idempotent.\nIf we use the StructureVisualizer, we\u0026rsquo;ll see a picture like this:\ngraph TD; Download_Task--\u0026gt; Print_Task; Print_Task; Summarizing For summarization, Griptape has many ready-made primitives, including TextSummaryTask. It takes a Summary Engine as input, through which you can configure summarization parameters, such as the model driver, prompt templates, and a certain Chunker, whose purpose we will discuss a little later. Let\u0026rsquo;s try to use this task:\nfrom griptape.tasks import TextSummaryTask from griptape.engines import PromptSummaryEngine from griptape.drivers.prompt.openai import OpenAiChatPromptDriver ... def process_url(url: str): ... download_task = CodeExecutionTask( on_run=load_page, input=url, id=\u0026#34;download_task\u0026#34;, child_ids=[\u0026#34;summary_task\u0026#34;] ) prompt_driver=OpenAiChatPromptDriver( model=\u0026#34;google/gemini-2.5-flash-preview-05-20\u0026#34;, base_url=\u0026#34;https://openrouter.ai/api/v1\u0026#34;, api_key=key, ) summary_task = TextSummaryTask( \u0026#34;Please summarize the following content in a concise manner. {{ parents_output_text }}\u0026#34;, summary_engine=PromptSummaryEngine( prompt_driver=prompt_driver ), id=\u0026#34;summary_task\u0026#34;, child_ids=[\u0026#34;print_task\u0026#34;], ) print_task = CodeExecutionTask(on_run=print_result, id=\u0026#34;print_task\u0026#34;) workflow = Workflow(tasks=[download_task, summary_task, print_task]) workflow.run() This code gives us the following output:\n$ uv run ./main.py https://docs.griptape.ai/stable/griptape-framework/structures/agents/ Output: Griptape Agents are a quick way to start using the platform. They take tools and input directly, which the agent uses to add a Prompt Task. The final output of the Agent can be accessed using the output attribute. An example demonstrates an Agent using a CalculatorTool to compute 13^7, successfully returning the result 62,748,517. Here I came across an interesting feature: for the summary_task to receive the output of download_task as its input, you must specify an instruction using the {parents_output_text} template as the input. Although the summarizer already has the necessary prompt inside, it does not extract the data from the output of the parent tasks on its own. We have to take care of that ourselves. Otherwise, everything seems quite straightforward.\ngraph TD; Download_Task--\u0026gt; Summary_Task; Summary_Task--\u0026gt; Print_Task; Print_Task Chunker While summarization worked great on small pages, trying to run it on a 55,000-word page caused the task to hang, regardless of the model used. Explicitly specifying a Chunker solved the problem:\nfrom griptape.chunkers import TextChunker ... summary_task = TextSummaryTask( \u0026#34;Please summarize the following content in a concise manner. {{ parents_output_text }}\u0026#34;, summary_engine=PromptSummaryEngine( prompt_driver=prompt_driver, chunker = TextChunker(max_tokens=16000) ), id=\u0026#34;summary_task\u0026#34;, child_ids=[\u0026#34;print_task\u0026#34;, \u0026#34;russian_translate_task\u0026#34;, \u0026#34;polish_translate_task\u0026#34;], ) Based on the logs, it seems the chunking in the summarizer is implemented iteratively, and in pseudocode, it can be expressed as follows:\nchunks = [chunk1, chunk2, ..., chunkN] summary = summarize(f\u0026#34;Summarize this: {chunk1}\u0026#34;) for chunk in chunks[1:]: summary = summarize(f\u0026#34;Update this summary {summary} with the following additional information: {chunk}\u0026#34;) return summary This approach causes information located closer to the end of the text to outweigh the information at the beginning, which is not always acceptable. Personally, I would prefer to summarize chunks via Map-Reduce:\nchunks = [chunk1, chunk2, ..., chunkN] summaries = [summarize(f\u0026#34;Summarize this: {chunk}\u0026#34;) for chunk in chunks] summary = summarize(f\u0026#34;Summarize these chunk summaries: {\u0026#39;\\n---\\n\u0026#39;.join(summaries)}\u0026#34;) return summary With this approach, all chunks are treated equally. It also allows for parallel summarization, although it requires one more request to the model.\nHowever, this is where the flexibility of the framework and its engines shines. Thanks to them, this algorithm can be implemented in your own custom engine, inherited from BaseSummaryEngine, and used almost seamlessly.\nBetter Together The next step is to translate this text into several languages. Of course, these tasks can be parallelized, and our Workflow provides excellent tools for this:\nfrom griptape.tasks import PromptTask ... def process_url(url: str): \u0026#34;\u0026#34;\u0026#34;Process the provided URL.\u0026#34;\u0026#34;\u0026#34; ... summary_task = TextSummaryTask( ... id=\u0026#34;summary_task\u0026#34;, child_ids=[\u0026#34;print_task\u0026#34;, \u0026#34;russian_translate_task\u0026#34;, \u0026#34;polish_translate_task\u0026#34;], ) russian_translate_task = PromptTask( \u0026#34;Please translate the following text to Russian: {{ parents_output_text }}\u0026#34;, prompt_driver=prompt_driver, id=\u0026#34;russian_translate_task\u0026#34;, child_ids=[\u0026#34;print_task\u0026#34;], ) polish_translate_task = PromptTask( \u0026#34;Please translate the following text to Polish: {{ parents_output_text }}\u0026#34;, prompt_driver=prompt_driver, id=\u0026#34;polish_translate_task\u0026#34;, child_ids=[\u0026#34;print_task\u0026#34;], ) print_task = CodeExecutionTask(on_run=print_result, id=\u0026#34;print_task\u0026#34;) workflow = Workflow( tasks=[ download_task, summary_task, russian_translate_task, polish_translate_task, print_task, ] ) That was surprisingly simple. The PromptTask is the most basic primitive used for direct requests to an LLM. The only interesting thing here is that we specified several children for summary_task, which allows multiple tasks to run in parallel.\nBut how can we verify that the tasks are actually running in parallel? The developers have thought of this too, providing support for hooks in the API. Specifically, the constructor of any task has on_before_run and on_after_run parameters, allowing you to add arbitrary pre- and post-processing. Let\u0026rsquo;s use them:\nfrom griptape.tasks import BaseTask from datetime import datetime ... def timestamp(task: BaseTask, action: str): print(f\u0026#34;task {task.id} {action} at {datetime.now().isoformat()}\u0026#34;) def process_url(url: str): ... polish_translate_task = PromptTask( \u0026#34;Please translate the following text to Polish: {{ parents_output_text }}\u0026#34;, prompt_driver=prompt_driver, on_before_run=lambda task: timestamp(task, \u0026#34;started\u0026#34;), on_after_run=lambda task: timestamp(task, \u0026#34;finished\u0026#34;), id=\u0026#34;polish_translate_task\u0026#34;, child_ids=[\u0026#34;print_task\u0026#34;], ) # And do the same for the other tasks We get the following result, which fully meets our expectations:\n$ uv run ./main.py https://docs.griptape.ai/stable/griptape-framework/structures/agents/ task summary_task started at 2025-06-05T22:03:02.272397 task summary_task finished at 2025-06-05T22:03:04.345067 task russian_translate_task started at 2025-06-05T22:03:04.347638 task polish_translate_task started at 2025-06-05T22:03:04.351531 task polish_translate_task finished at 2025-06-05T22:03:05.702819 task russian_translate_task finished at 2025-06-05T22:03:05.947364 Output: Agents in Griptape offer a quick start, directly processing tools and input to generate a Prompt Task. The final output is accessible via the `output` attribute. An example demonstrates an Agent using a `CalculatorTool` to compute 13^7, showing the input, tool action, and the resulting output. Output: Вот перевод текста на русский язык: Агенты в Griptape предлагают быстрый старт, напрямую обрабатывая инструменты и входные данные для генерации задачи (Prompt Task). Конечный результат доступен через атрибут `output`. Пример демонстрирует Агента, использующего `CalculatorTool` для вычисления 13^7, показывая входные данные, действие инструмента и полученный результат. Output: Oto tłumaczenie tekstu na język polski: Agenci w Griptape oferują szybki start, bezpośrednio przetwarzając narzędzia i dane wejściowe w celu wygenerowania Zadania Monitu (Prompt Task). Ostateczny wynik jest dostępny poprzez atrybut `output`. Przykład demonstruje Agenta używającego `CalculatorTool` do obliczenia 13^7, pokazując dane wejściowe, działanie narzędzia i wynik końcowy. The logs clearly show that the Russian and Polish translations are running in parallel.\nAnd for good measure, here\u0026rsquo;s the current graph:\ngraph TD; Download_Task--\u0026gt; Summary_Task; Summary_Task--\u0026gt; Print_Task \u0026amp; Russian_Translate_Task \u0026amp; Polish_Translate_Task; Russian_Translate_Task--\u0026gt; Print_Task; Polish_Translate_Task--\u0026gt; Print_Task; Print_Task; Finishing Up The rest is essentially a matter of technique, so I\u0026rsquo;ll just provide the complete code for the program below:\nimport argparse import dotenv import os from griptape.structures import Workflow from griptape.tasks import CodeExecutionTask from griptape.loaders import WebLoader from griptape.utils import StructureVisualizer from griptape.tasks import TextSummaryTask from griptape.engines import PromptSummaryEngine from griptape.drivers.prompt.openai import OpenAiChatPromptDriver from griptape.tasks import PromptTask, BaseTask from griptape.chunkers import TextChunker from griptape.artifacts import TextArtifact from datetime import datetime import logging logging.getLogger(\u0026#34;griptape\u0026#34;).setLevel(logging.WARNING) dotenv.load_dotenv() def load_page(task: CodeExecutionTask) -\u0026gt; str: \u0026#34;\u0026#34;\u0026#34;Load the content of the given URL.\u0026#34;\u0026#34;\u0026#34; return WebLoader().load(task.input.value) def combine_result(task: CodeExecutionTask) -\u0026gt; TextArtifact: \u0026#34;\u0026#34;\u0026#34;Combine results from parent tasks.\u0026#34;\u0026#34;\u0026#34; result = \u0026#34;\u0026#34; for parent in task.parents: if parent.output.value: result += f\u0026#34;{parent.output.value}\\n\\n\u0026#34; return TextArtifact(result) def send_to_telegram(task: CodeExecutionTask) -\u0026gt; None: \u0026#34;\u0026#34;\u0026#34;Send the result to Telegram.\u0026#34;\u0026#34;\u0026#34; # Placeholder for sending to Telegram logic print(f\u0026#34;Sending to Telegram: {task.parents[0].output.value}\u0026#34;) def timestamp(task: BaseTask, action: str): print(f\u0026#34;task {task.id} {action} at {datetime.now().isoformat()}\u0026#34;) def process_url(url: str): \u0026#34;\u0026#34;\u0026#34;Process the provided URL.\u0026#34;\u0026#34;\u0026#34; key = os.environ.get(\u0026#34;OPENROUTER_API_KEY\u0026#34;, \u0026#34;\u0026#34;) if not key: raise ValueError(\u0026#34;OPENROUTER_API_KEY is not set in the environment variables.\u0026#34;) prompt_driver=OpenAiChatPromptDriver( model=\u0026#34;google/gemini-2.5-flash-preview-05-20\u0026#34;, base_url=\u0026#34;https://openrouter.ai/api/v1\u0026#34;, api_key=key, ) download_task = CodeExecutionTask( on_run=load_page, input=url, id=\u0026#34;download_task\u0026#34;, child_ids=[\u0026#34;summary_task\u0026#34;] ) summary_task = TextSummaryTask( \u0026#34;Please summarize the following content in a concise manner. {{ parents_output_text }}\u0026#34;, summary_engine=PromptSummaryEngine( prompt_driver=prompt_driver, chunker = TextChunker(max_tokens=16000) ), on_before_run=lambda task: timestamp(task, \u0026#34;started\u0026#34;), on_after_run=lambda task: timestamp(task, \u0026#34;finished\u0026#34;), id=\u0026#34;summary_task\u0026#34;, child_ids=[\u0026#34;combine_task\u0026#34;, \u0026#34;russian_translate_task\u0026#34;, \u0026#34;polish_translate_task\u0026#34;], ) russian_translate_task = PromptTask( \u0026#34;Please translate the following text to Russian: {{ parents_output_text }}\u0026#34;, prompt_driver=prompt_driver, on_before_run=lambda task: timestamp(task, \u0026#34;started\u0026#34;), on_after_run=lambda task: timestamp(task, \u0026#34;finished\u0026#34;), id=\u0026#34;russian_translate_task\u0026#34;, child_ids=[\u0026#34;combine_task\u0026#34;], ) polish_translate_task = PromptTask( \u0026#34;Please translate the following text to Polish: {{ parents_output_text }}\u0026#34;, prompt_driver=prompt_driver, on_before_run=lambda task: timestamp(task, \u0026#34;started\u0026#34;), on_after_run=lambda task: timestamp(task, \u0026#34;finished\u0026#34;), id=\u0026#34;polish_translate_task\u0026#34;, child_ids=[\u0026#34;combine_task\u0026#34;], ) combine_task = CodeExecutionTask(on_run=combine_result, id=\u0026#34;combine_task\u0026#34;, child_ids=[\u0026#34;send_task\u0026#34;]) send_task = CodeExecutionTask( on_run=send_to_telegram, id=\u0026#34;send_task\u0026#34;, ) workflow = Workflow( tasks=[ download_task, summary_task, russian_translate_task, polish_translate_task, combine_task, send_task ] ) workflow.run() print(StructureVisualizer(workflow).to_url()) print(\u0026#34;Workflow completed successfully.\u0026#34;) def main(): parser = argparse.ArgumentParser(description=\u0026#34;Process URLs for link blog\u0026#34;) parser.add_argument(\u0026#34;url\u0026#34;, type=str, help=\u0026#34;URL to process\u0026#34;) args = parser.parse_args() process_url(args.url) if __name__ == \u0026#34;__main__\u0026#34;: main() And the graph:\ngraph TD; Download_Task--\u0026gt; Summary_Task; Summary_Task--\u0026gt; Combine_Task \u0026amp; Russian_Translate_Task \u0026amp; Polish_Translate_Task; Russian_Translate_Task--\u0026gt; Combine_Task; Polish_Translate_Task--\u0026gt; Combine_Task; Combine_Task--\u0026gt; Send_Task; Send_Task; I find the code to be quite simple, easy to read, and flexible enough for modification and reuse. Obviously, in real-world applications, all this simplicity will be diluted with error handling, logging, and so on. But that\u0026rsquo;s a topic for a separate discussion.\nAdditionally, the Workflow API provides two more styles for task composition that don\u0026rsquo;t require specifying parents and children during task creation:\nImperative, where we can use add_parent and add_child functions. The so-called bit-shift, where parents and children can be linked like this: task1 \u0026gt;\u0026gt; task2 \u0026gt;\u0026gt; [task3, task4] Both of these styles make it easy to assemble different graphs from ready-made primitives, although bit-shift feels more like a separate DSL than standard Python.\nWe\u0026rsquo;re Done In conclusion, I\u0026rsquo;d like to say that I\u0026rsquo;m enjoying the framework at this stage. It is quite logical, flexible, and pleasant to use, although it is not without some rough edges.\nI also found the documentation to be quite good, although it lacks a description of exactly how some of the engines work. For now, to get this information, you have to dig through the logs or the code.\nIn this post we\u0026rsquo;ve figured out the basic functionality. Next time, we\u0026rsquo;ll look at what primitives the framework provides for building RAGs.\nwe want to make sure we\u0026rsquo;re not painfully embarrassed by what we\u0026rsquo;ve published, right?\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nwhich I highly respect and hope to write a separate post about someday\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://meshrefine.com/en/posts/griptape-2/","summary":"\u003cp\u003eIn the \u003ca href=\"/en/posts/griptape-1/\"\u003eprevious post\u003c/a\u003e, I broke down the basic concepts of the \u003ca href=\"https://www.griptape.ai\"\u003eGriptape\u003c/a\u003e AI framework, and now it\u0026rsquo;s time to put them into practice. We\u0026rsquo;ll try to use them to develop a small application that helps run a link-blog on Telegram.\u003c/p\u003e\n\u003cp\u003eThe application will receive a URL, download its content, run it through an LLM to generate a summary, translate that summary into a couple of other languages, combine everything, and publish it to Telegram via a bot. The general flow can be seen in the diagram below:\u003c/p\u003e","title":"Griptape, Part 2: Building Graphs"},{"content":"What on Earth is Codex? Good question, right? The thing is, until recently, OpenAI had a model called Codex, which was used as the foundation for autocompletion in GitHub Copilot. Then, OpenAI released a console agent for development, which they named, so no one would get confused, Codex. 1 Everyone had a laugh at OpenAI\u0026rsquo;s naming skills 2, and life went on. Until the fateful day when a tweet like this appeared from Sam Altman:\nI was intrigued. Firstly, by the hype-building phrase \u0026ldquo;low-key research preview\u0026rdquo; 3, and secondly, by that very name. However, I wasn\u0026rsquo;t disappointed:\nWith this post, the family of entities named Codex was expanded by two new members:\nAn o3 model variant, specifically fine-tuned for programming, named codex-1. A cloud-based agent capable of autonomously performing several different tasks on a GitHub repository 4. It\u0026rsquo;s the latter we\u0026rsquo;ll be talking about today.\nHow does it work? The sequence of actions required to use the agent is quite simple:\nStep 1. We go to https://chatgpt.com/codex. There, we see a list of tasks the agent is executing/has executed, and an input window where we can select the repository we want to work with and the branch. There are two buttons - \u0026ldquo;Ask,\u0026rdquo; which analyzes the code before automatically suggesting specific sub-tasks, and \u0026ldquo;Code,\u0026rdquo; which will directly edit the code.\nStep 2. To add a new repository, we go to the Environments menu and create a new environment there. We can specify the repo itself, configure the workspace settings, and test it out.\nStep 3. We return to the main screen and write our request.\nStep 4, the most interesting part. The environment deploys an Ubuntu-based container, installs the necessary packages, applies custom environment settings, and clones the repository inside. After this, the internet is disconnected 5 (not always anymore, which is what this post is about), and the agent begins its work. It does this long and diligently, as the scheduler is from o3, and it\u0026rsquo;s a good scheduler. This process can be observed in the Logs window:\nStep 5. After several minutes of work, we are presented with a diff, which we can review and either create a pull-request on GitHub or steer the process back on the right track and repeat the iteration.\nSteps 3-5 can be run in parallel, allowing you to go about your business while the agent works. The diffs it produces are very compact and pleasant to read, which increases confidence in the correctness of the result.\nInternet Access If you look closely at the process, you can see a certain limitation. The lack of internet access during execution breaks many build and testing processes, which limits the agent\u0026rsquo;s capabilities. This usually leads to correct results with the note \u0026ldquo;Running tests: failed.\u0026rdquo; This happens for various reasons, but primarily because it\u0026rsquo;s often necessary to download various libraries during the build. Of course, this can be circumvented during the environment setup stage when access is still available, but this process is far from always trivial.\nThis limitation was quite painful, which is why OpenAI recently announced that internet access can now be left on for the model. Of course, letting the model run wild in the pasture onto the internet without any restrictions is dangerous, so we were given a choice of access modes.\nFirstly, we can leave everything as is and not allow internet access. Secondly, we can grant access, but to a limited number of resources 6. Additionally, we can allow the model to perform only read operations. The most daring and brave are allowed to give the agent full and unrestricted access and hope for the best.\nSecurity Why hope? Because direct access to unverified resources threatens a whole host of problems, about which OpenAI persistently warns.\nIn this warning, various concerns are all jumbled together, mentioning an attack and its consequences in the same list:\nPrompt injection: The model downloads a webpage, sees text that looks like an instruction, and happily executes it. The instruction might very well ask it to\u0026hellip; Exfiltration of code or secrets: \u0026hellip;upload the repository contents and secrets to a third-party resource; Inclusion of malware or vulnerabilities: \u0026hellip;or use a malicious library instead of a legitimate one. Use content with license restriction: However, the model might decide on its own that the code it found in a repository under a GPL license is the best thing to add to our repository, which could lead to certain legal problems. Therefore, it is strictly recommended to limit the model\u0026rsquo;s access only to trusted resources, and only when necessary.\nA Small Example and Conclusion Of course, internet access opens up a lot of interesting possibilities beyond simplifying builds. For example, the first thing I asked the model to do was to check this website for potential SEO issues. The result can be seen below.\nTo sum up, I can say that I generally like the tool. The small size of the diffs and the ability to run multiple tasks in parallel allow for small refactorings and improvements on a fire-and-forget basis, without having to dive deep into the context or get distracted from other tasks. This saves a lot of time and reduces cognitive load.\nIt will be interesting to see what competitors offer in response. 7\nWith the CLI suffix, although it\u0026rsquo;s referred to almost everywhere simply as Codex.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nAfter all, compared to their model naming, such a mix-up is child\u0026rsquo;s play.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nSam is a master of mutually exclusive statements.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nOthers are not (yet?) supported.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nFor security reasons, which we\u0026rsquo;ll discuss further.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nIn this case, the list of resources with various repositories is already pre-defined.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nA few days later at Google I/O 2025, Jules was introduced, working on a very similar principle, but due to the large size of the diffs generated by Gemini 2.5 Pro, it\u0026rsquo;s significantly harder to use.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://meshrefine.com/en/posts/openai-codex-internet-access/","summary":"\u003ch2 id=\"what-on-earth-is-codex\"\u003eWhat on Earth is Codex?\u003c/h2\u003e\n\u003cp\u003eGood question, right? The thing is, until recently, OpenAI had a model called Codex, which was used as the foundation for autocompletion in GitHub Copilot. Then, OpenAI released a console agent for development, which they named, so no one would get confused, \u003ca href=\"https://github.com/openai/codex\"\u003eCodex\u003c/a\u003e. \u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e Everyone had a laugh at OpenAI\u0026rsquo;s naming skills \u003csup id=\"fnref:2\"\u003e\u003ca href=\"#fn:2\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e2\u003c/a\u003e\u003c/sup\u003e, and life went on. Until the fateful day when a tweet like this appeared from Sam Altman:\u003c/p\u003e","title":"OpenAI Codex Gains Internet Access: First Impressions"},{"content":" 1. Introduction: The Imperative of AI Crawler Management 2. Understanding robots.txt: The Foundation of Crawler Instruction 2.1. Core Syntax and Directives 2.2. File Placement and Formatting 2.3. How Crawlers Interpret robots.txt 2.4. Testing Your robots.txt 3. Identifying Key Crawler Types: AI Agents vs. Search Engine Bots 3.1. Distinguishing Characteristics 3.2. Categories of AI Crawlers and Their User Agents 3.2.1. AI Crawlers for Model Training 3.2.2. AI Crawlers for Live Retrieval and Search Assistance 3.3. Standard Search Engine Crawlers (to be Allowed) 3.4. Table of Prominent AI Crawler User Agents 4. Strategically Blocking AI Crawlers with robots.txt 4.1. Targeting Specific AI User Agents 4.2. Applying Rules to Specific Pages or Directories 4.3. Ensuring Search Engines Are Not Blocked from Specific Pages 4.4. The Challenge of \u0026ldquo;All Possible AI\u0026rdquo; 5. Advanced Methods for Granular AI Crawler Control 5.1. HTML Meta Tags (Page-Level Control) 5.2. HTTP X-Robots-Tag Headers (Server-Level Page Control) 5.3. Server-Side Blocking 5.4. Web Application Firewalls (WAFs) and Content Delivery Networks (CDNs) 5.5. Table: Comparison of AI Crawler Control Mechanisms 6. Limitations and Best Practices 6.1. robots.txt is a Directive, Not an Enforcement Mechanism 6.2. Importance of Regular Review and Updates 6.3. Testing robots.txt Changes 6.4. Avoiding Common Pitfalls 6.5. Log File Analysis 7. Conclusion: Implementing a Robust, Layered AI Crawler Defense 1. Introduction: The Imperative of AI Crawler Management The proliferation of Artificial Intelligence (AI) has introduced a new class of web crawlers designed to gather vast quantities of data for training Large Language Models (LLMs) and powering AI-driven applications. While these advancements offer significant potential, website operators often require precise control over which content AI crawlers can access, particularly to protect intellectual property, sensitive information, or manage server resources. Simultaneously, maintaining visibility and crawlability for traditional search engine bots like Googlebot and Bingbot remains paramount for organic search performance.\nThis report provides an expert-level guide on utilizing the robots.txt protocol and other webmaster tools to selectively prevent AI crawlers from parsing specific pages or sections of a website, without impeding the access of legitimate search engine crawlers. It delves into the intricacies of robots.txt syntax, strategies for identifying and targeting AI bots, the application of advanced control mechanisms beyond robots.txt, and best practices for maintaining an effective and evolving crawler management strategy. The objective is to equip webmasters with the knowledge to implement robust and nuanced control over how various automated agents interact with their web properties.\n2. Understanding robots.txt: The Foundation of Crawler Instruction The Robots Exclusion Protocol (REP), commonly implemented via the robots.txt file, serves as the primary method for webmasters to communicate their crawling preferences to web robots. While not a security mechanism, it is a widely respected standard by reputable crawlers, including those operated by major search engines and many AI companies.\n2.1. Core Syntax and Directives A robots.txt file is a simple plain text file consisting of one or more rules, or groups of directives. Each group typically begins with a User-agent line, specifying the crawler(s) the rules apply to, followed by Disallow or Allow directives.\nUser-agent: This directive specifies the name (token) of the web crawler to which the subsequent rules in the group apply. A wildcard, *, can be used to indicate all crawlers, unless a more specific user-agent rule matches the crawler.\nExample: User-agent: GPTBot targets OpenAI\u0026rsquo;s GPTBot. Example: User-agent: * targets all bots that do not have a more specific rule set. Disallow: This directive instructs the specified user-agent not to crawl particular paths. The path should be the part of the URL that comes after the domain name, starting with a /.\nExample: Disallow: /private-data/ blocks access to all content within the /private-data/ directory. Example: Disallow: /specific-page.html blocks access to that single HTML file. A Disallow: / directive for a specific user-agent blocks it from crawling the entire site. Allow: This directive explicitly permits the specified user-agent to crawl a path, even if it falls within a disallowed directory. This is particularly useful for allowing access to a specific file or subdirectory within an otherwise disallowed section. Googlebot and Bingbot support this directive.\nExample:\nUser-agent: Googlebot Disallow: /reports/ Allow: /reports/public-summary.pdf This would block Googlebot from the /reports/ directory but allow it to access public-summary.pdf within that directory.\nSitemap: This directive, while not part of the original REP, is widely supported and used to specify the location of XML sitemap(s). It helps crawlers discover all relevant URLs on a site. Multiple Sitemap directives can be included.\nExample: Sitemap: https://www.example.com/sitemap.xml Comments (#): Lines beginning with a # character are treated as comments and are ignored by crawlers. They are useful for adding human-readable notes and explanations within the robots.txt file, which is essential for maintaining clarity, especially when managing numerous rules for various AI bots.\nThe straightforward nature of robots.txt, being a simple text file with a limited set of commands, is a significant advantage, making it accessible for webmasters of all skill levels to implement basic crawler instructions. However, this simplicity is also a source of its primary limitation: its effectiveness hinges entirely on the voluntary compliance of web crawlers. Bots designed with malicious intent, or those operated by entities that choose not to adhere to the Robots Exclusion Protocol, will simply ignore the directives. Consequently, while robots.txt serves as an important first line of communication for expressing crawling preferences, it should not be considered a security measure. For content that requires robust protection from unauthorized access, methods such as server-side authentication, IP address blocking, or Web Application Firewalls (WAFs) are necessary complements, forming part of a layered defense strategy.\n2.2. File Placement and Formatting For robots.txt to be effective, it must adhere to specific placement and formatting rules:\nThe file must be named exactly robots.txt, in lowercase. It must be located at the root of the website\u0026rsquo;s host. For a site https://www.example.com, the robots.txt file must be accessible at https://www.example.com/robots.txt. It cannot be placed in a subdirectory. A website can only have one robots.txt file. If multiple files were allowed, it would create ambiguity for crawlers. The file must be a UTF-8 encoded text file. ASCII is a subset of UTF-8 and is also acceptable. Using other encodings may lead to characters being misinterpreted, potentially invalidating rules. 2.3. How Crawlers Interpret robots.txt Crawlers that respect the REP typically follow a standard procedure:\nBefore crawling any other URLs on a host, a crawler will attempt to fetch the robots.txt file. Rules are organized into groups, and crawlers process these groups from top to bottom. A user agent will attempt to find the group of rules that most specifically matches its user-agent string. It will obey the rules in the first such specific group it finds. All other groups are ignored by that user agent. If multiple groups specify the same user agent, compliant crawlers will combine the directives from these groups into a single conceptual group before processing. Implicit Allowance: A crucial aspect of the REP is that any URL not explicitly disallowed by a matching Disallow directive is implicitly allowed for crawling. This principle is fundamental to the strategy of allowing search engines by default while selectively blocking AI crawlers. The \u0026ldquo;first match, most specific group\u0026rdquo; rule has significant implications for the structure and ordering of directives within a robots.txt file. When crafting rules to differentiate between AI crawlers and search engine bots, particularly for specific paths, the order can be critical. For instance, if a general rule for User-agent: * disallows a directory, but a subsequent, more specific rule for User-agent: Googlebot allows access to that same directory, Googlebot will follow its specific rule. However, if the rules were ordered differently, or if an AI bot\u0026rsquo;s user-agent string inadvertently matched a broad rule intended for another purpose, unintended blocking or allowing could occur. This underscores the necessity for careful planning and testing, especially as the list of AI bots to manage grows. More specific user-agents (like individual AI bot tokens) should generally be defined with their rules before more general ones (*) if there\u0026rsquo;s a potential for conflicting path directives.\n2.4. Testing Your robots.txt After creating or modifying a robots.txt file, it is essential to test its validity and ensure it behaves as expected:\nPublic Accessibility: Verify that the file is publicly accessible by navigating to its URL (e.g., https://www.example.com/robots.txt) in a private browsing window. You should see the plain text content of your file. Syntax and Logic Testing: Tools like the robots.txt Tester in Google Search Console allow webmasters to validate their file, check if specific URLs are blocked or allowed for Google\u0026rsquo;s crawlers, and identify syntax errors. Similar tools may be available from other search engine providers or third-party SEO platforms. 3. Identifying Key Crawler Types: AI Agents vs. Search Engine Bots Effectively managing crawler access requires distinguishing between different types of bots, primarily traditional search engine crawlers and the newer generation of AI agents.\n3.1. Distinguishing Characteristics The primary difference lies in their purpose.\nSearch engine bots (e.g., Googlebot, Bingbot) crawl the web to discover, index, and rank content for inclusion in search engine results pages, with the goal of making information findable by users. AI crawlers gather data for a broader range of AI-related tasks. This includes collecting massive datasets of text, images, and code to train LLMs (e.g., GPTBot, Google-Extended, ClaudeBot), or fetching real-time information from the web to provide up-to-date answers in AI chat interfaces or search-like applications (e.g., ChatGPT-User, Perplexity-User). Each bot identifies itself using a specific user-agent token in its HTTP requests. Recognizing these tokens is the cornerstone of targeting them with robots.txt directives.\n3.2. Categories of AI Crawlers and Their User Agents AI crawlers can be broadly categorized based on their primary function:\n3.2.1. AI Crawlers for Model Training These bots are focused on amassing data to build and refine the foundational knowledge of AI models. Examples include:\nGPTBot: OpenAI\u0026rsquo;s crawler for training generative AI models. Google-Extended: Google\u0026rsquo;s user agent for data collection to improve Gemini, Vertex AI, and future generative models. Blocking this does not affect Google Search ranking or inclusion. ClaudeBot: Anthropic\u0026rsquo;s primary web crawler for training its LLMs, such as Claude. anthropic-ai: Another user agent associated with Anthropic, potentially for specific development purposes or a legacy bot. CCBot: Common Crawl\u0026rsquo;s bot, which archives vast swathes of the web. This data is publicly available and frequently used to train AI models by various organizations. Amazonbot: Amazon\u0026rsquo;s crawler, used for services like Alexa and likely for training Amazon\u0026rsquo;s LLMs. Bytespider: ByteDance\u0026rsquo;s (parent company of TikTok) crawler, likely used for LLM training. It has been reported to sometimes ignore robots.txt directives. Meta-ExternalAgent (formerly FacebookBot): Meta\u0026rsquo;s crawler for AI model training and other services. cohere-ai: Cohere\u0026rsquo;s bot for collecting text samples to refine its language models. Applebot-Extended: Apple\u0026rsquo;s bot used to determine how data crawled by Applebot can be used for Apple\u0026rsquo;s foundation models. GoogleOther: Used by Google for internal research and development, which may include model training. 3.2.2. AI Crawlers for Live Retrieval and Search Assistance These bots retrieve current information from websites to answer user queries in real-time within AI applications.\nChatGPT-User: OpenAI\u0026rsquo;s bot that facilitates web browsing within ChatGPT, enabling it to access live information. PerplexityBot / Perplexity-User: Perplexity AI uses PerplexityBot to build and maintain its own search index (explicitly stated as not for AI model training). Perplexity-User supports live user queries within Perplexity and is documented to generally ignore robots.txt rules because the fetch is user-initiated. OAI-SearchBot: OpenAI\u0026rsquo;s crawler used to create an index for its SearchGPT product. DuckAssistBot: DuckDuckGo\u0026rsquo;s bot for collecting data to deliver AI-backed answers. The differentiation between AI crawlers for \u0026ldquo;model training\u0026rdquo; and those for \u0026ldquo;live retrieval\u0026rdquo; is an important nuance. While the current objective may be to block specific pages from all AI, some website operators might in the future consider a more granular approach. For instance, they might choose to block bots that train models on their content to protect intellectual property 13, while simultaneously allowing live retrieval bots if they perceive a benefit in their content being accurately cited and surfaced in AI-assisted search results. However, this nuanced strategy is complicated by the behavior of certain bots, like Perplexity-User, which explicitly state they ignore robots.txt for user-initiated fetches. This indicates that for bots bypassing robots.txt, more assertive control methods such as IP blocking or WAF rules would be necessary to enforce such distinctions.\nThe emergence of distinct AI-specific user-agent tokens, such as Google-Extended separate from the traditional Googlebot 14, and OpenAI\u0026rsquo;s differentiation between GPTBot and ChatGPT-User 6, signals a recognition by major technology companies of webmasters\u0026rsquo; desire for differentiated control over data usage. This trend may eventually lead to more standardized protocols for declaring AI interaction policies. However, in the current landscape, it translates to an increased number of user-agent tokens that webmasters must identify, track, and manage within their robots.txt files. If a company does not provide a distinct token for its AI-related crawling activities, its primary search bot might be performing dual roles, making it challenging to restrict AI data usage without potentially impacting search engine visibility.\n3.3. Standard Search Engine Crawlers (to be Allowed) For the purpose of this report, it is crucial to ensure that directives aimed at AI crawlers do not inadvertently block standard search engine bots. Key search engine user agents include:\nGooglebot: Google\u0026rsquo;s main crawler for web search. Bingbot: Microsoft\u0026rsquo;s crawler for Bing search. DuckDuckBot: DuckDuckGo\u0026rsquo;s web crawler. Slurp: Yahoo\u0026rsquo;s historic crawler (less prevalent but may still be encountered). YandexBot: Yandex\u0026rsquo;s crawler. Applebot: Apple\u0026rsquo;s crawler for Siri and Spotlight suggestions. (Note the distinction from Applebot-Extended used for foundation models). 3.4. Table of Prominent AI Crawler User Agents The following table summarizes key AI crawler user agents relevant for robots.txt management. The \u0026ldquo;User-Agent Token\u0026rdquo; is the string to use in the User-agent: line in robots.txt.\nTable 1: Prominent AI Crawler User Agents for robots.txt\nAI Company robots.txt User-Agent Token Full User-Agent String (Example) Primary Purpose Respects robots.txt? OpenAI GPTBot Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.1; +https://openai.com/gptbot) Model Training Yes OpenAI ChatGPT-User Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot) Live Retrieval for ChatGPT Yes OpenAI OAI-SearchBot Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot) Indexing for OpenAI Search Yes Google Google-Extended Mozilla/5.0 (compatible; Google-Extended/1.0; +http://www.google.com/bot.html) Model Training (Gemini, Vertex AI) Yes Anthropic ClaudeBot Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ClaudeBot/1.0; +claudebot@anthropic.com) Model Training Yes (Assumed) Anthropic anthropic-ai Mozilla/5.0 (compatible; anthropic-ai/1.0; +http://www.anthropic.com/bot.html) Model Training (Potentially legacy) Yes (Assumed) Common Crawl CCBot Mozilla/5.0 (compatible; CCBot/1.0; +http://www.commoncrawl.org/bot.html) Open Web Data Archiving (used for AI training) Yes Perplexity AI PerplexityBot Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot) Indexing for Perplexity Search (not for training) Yes Perplexity AI Perplexity-User Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user) Live Retrieval for Perplexity Ignores robots.txt ByteDance Bytespider Mozilla/5.0 (compatible; Bytespider/1.0; +http://www.bytedance.com/bot.html) Model Training (TikTok) Often Ignores Meta Meta-ExternalAgent Mozilla/5.0 (compatible; meta-externalagent/1.1; +https://developers.facebook.com/docs/sharing/webmasters/crawler) Model Training Yes (Assumed) Apple Applebot-Extended Mozilla/5.0 (compatible; Applebot-Extended/1.0; +http://www.apple.com/bot.html) Training Apple\u0026rsquo;s foundation models Yes (Assumed) Cohere cohere-ai Mozilla/5.0 (compatible; cohere-ai/1.0; +http://www.cohere.ai/bot.html) Model Training Yes (Assumed) Google GoogleOther GoogleOther Internal R\u0026amp;D, potentially model training Yes (Assumed) Note: \u0026ldquo;Yes (Assumed)\u0026rdquo; indicates that while not explicitly stated for every bot in the provided materials, reputable AI companies generally claim to respect robots.txt. However, verification through log analysis is always recommended.\n4. Strategically Blocking AI Crawlers with robots.txt The core strategy for preventing AI crawlers from accessing specific pages, while allowing search engines, involves using targeted Disallow directives for known AI user-agent tokens.\n4.1. Targeting Specific AI User Agents For each AI crawler identified (as per Table 1 or ongoing research), a distinct User-agent group should be created in the robots.txt file. Within each group, Disallow directives will specify the paths the bot is not permitted to crawl.\nExample structure:\nUser-agent: GPTBot Disallow: /confidential-research/ Disallow: /private-data/ Disallow: /specific-page-for-ai-block.html User-agent: ClaudeBot Disallow: /confidential-research/ # Same paths as GPTBot Disallow: /private-data/ Disallow: /specific-page-for-ai-block.html #...and so on for all other AI crawlers to be blocked from these paths. 4.2. Applying Rules to Specific Pages or Directories The Disallow directive is path-specific:\nTo block an entire directory and all its contents: Disallow: /directory-name/. Ensure the trailing slash is used if you intend to block the directory itself and everything under it. To block a single file: Disallow: /path/to/specific-file.html. All paths must start with a / and represent the path from the site root. Paths are generally case-sensitive. 4.3. Ensuring Search Engines Are Not Blocked from Specific Pages The primary goal is to block AI crawlers from specific pages/directories, not to block search engines from those same locations. This is achieved through the specificity of robots.txt rules:\nImplicit Allowance: Because search engine bots like Googlebot or Bingbot will not match the User-agent tokens specified for AI crawlers (e.g., GPTBot, ClaudeBot), the Disallow rules under those AI-specific groups will not apply to them. If there are no other rules in robots.txt that would disallow Googlebot (or other search engines) from accessing /confidential-research/, then Googlebot is implicitly allowed to crawl it. No Explicit Allow Needed (for this specific goal): For the user\u0026rsquo;s stated goal, explicit Allow: directives for search engines for these specific paths are generally not necessary. The absence of a matching Disallow rule for their user-agent is sufficient for them to crawl those paths. An explicit Allow would only be needed if, for example, a very broad rule like User-agent: * Disallow: / was in place (which is not recommended for this scenario as it would block all search engines from everything by default), or if a search engine needed access to a sub-path of a directory that was disallowed for that same search engine. The strategy of individually listing Disallow rules for each AI crawler can lead to a lengthy robots.txt file, especially if many distinct paths are being protected from numerous bots. While Google processes robots.txt files up to 500KB in size, which is substantial, extremely verbose files could theoretically approach this limit. This consideration might encourage webmasters to be as concise as possible or to explore server-side methods if the robots.txt file becomes unwieldy. However, the standard robots.txt protocol does not offer a native way to group multiple distinct User-agent tokens to share a single block of Disallow directives; each User-agent line typically starts a new rule group, or multiple User-agent lines at the beginning of a group apply to all directives within that group until the next User-agent line or the end of the file. Thus, repeating Disallow rules for each AI agent is the common and correct approach.\n4.4. The Challenge of \u0026ldquo;All Possible AI\u0026rdquo; It is practically impossible to block \u0026ldquo;all possible\u0026rdquo; AI crawlers, especially future or unknown ones, using robots.txt alone. The method relies on knowing the specific user-agent tokens these crawlers use. New AI bots are continually emerging, and some may not publicly document their user-agent strings or may attempt to masquerade as common browser user agents to evade detection.\nThe most effective robots.txt-based strategy is to:\nBe comprehensive with the list of known AI crawlers (referencing resources like Dark Visitors 13 or industry lists). Regularly review and update the robots.txt file as new AI crawlers are identified or as existing ones change their tokens. Avoid using overly broad User-agent: * Disallow: /some/path/ rules if the intent is only to block AI, as this could inadvertently block new, legitimate, non-search-engine services or even misconfigured search bots. The query specifically requires search engines not to be blocked from these paths. The very act of webmasters meticulously curating robots.txt files to block specific AI crawlers sends a collective signal to AI development companies. It indicates that their crawling activities are being actively monitored and that there is a clear demand from the web community for more transparent, controllable, and respectful AI data collection mechanisms. This widespread adoption of AI-specific robots.txt rules, as evidenced by statistics like 22% of top websites blocking GPTBot and CCBot 13, can contribute to a feedback loop. This may incentivize AI companies to ensure stricter adherence to robots.txt, provide clearer documentation for their bots, offer dedicated user-agents for different functions, and potentially participate more actively in the development of new web standards for granular control over data usage in AI contexts.\n5. Advanced Methods for Granular AI Crawler Control While robots.txt is the foundational layer, several other tools and techniques can provide more granular or forceful control over AI crawler access, especially for bots that may not fully respect robots.txt or when page-specific directives are desired.\n5.1. HTML Meta Tags (Page-Level Control) HTML meta tags, placed within the \u0026lt;head\u0026gt; section of an individual HTML page, can signal preferences to bots that are programmed to recognize them.\nnoai: This proposed directive aims to tell AI bots not to use the page\u0026rsquo;s content for training purposes. Example: \u0026lt;meta name=\u0026quot;robots\u0026quot; content=\u0026quot;noai\u0026quot;\u0026gt; It can also be targeted to specific bots: \u0026lt;meta name=\u0026quot;googlebot\u0026quot; content=\u0026quot;noai\u0026quot;\u0026gt; or \u0026lt;meta name=\u0026quot;gptbot\u0026quot; content=\u0026quot;noindex\u0026quot;\u0026gt; (though noindex for gptbot would prevent indexing by its search functions, noai would be more specific for training). noimageai: This proposed directive aims to prevent AI from using images on the page for model training. Example: \u0026lt;meta name=\u0026quot;robots\u0026quot; content=\u0026quot;noai, noimageai\u0026quot;\u0026gt; noml (No Machine Learning): A newer proposal, functionally similar to noai, intended to prevent content from being used for any machine learning purposes. Example: \u0026lt;meta name=\u0026quot;robots\u0026quot; content=\u0026quot;noml\u0026quot;\u0026gt; Effectiveness and Adoption: These meta tags are currently informal and not universally standardized or respected by all AI crawlers. However, their adoption is growing, and they can serve as an additional layer of instruction for compliant bots. They offer page-specific granularity, which robots.txt (being path-based) does not provide as directly for individual file content usage.\nThe emergence of these page-level tags like noai and noml, even if not yet universally adopted, points to a significant trend: a push from the web community for more standardized, machine-readable methods to express data usage preferences specifically for AI. The robots.txt protocol primarily dictates whether a path can be crawled or not; it doesn\u0026rsquo;t inherently convey instructions about how crawled content may be used. These new meta tags attempt to bridge this gap by directly addressing the \u0026ldquo;use for AI training\u0026rdquo; concern at a granular, page-by-page level. This reflects a broader desire for more nuanced control over data usage rather than just data access, essentially signaling: \u0026ldquo;You may crawl this page for search indexing, but you may not use its content to train your AI model.\u0026rdquo; If widely adopted by both websites and AI crawlers, these tags could form a more explicit framework for consent in AI data consumption.\n5.2. HTTP X-Robots-Tag Headers (Server-Level Page Control) Directives such as noai and noimageai can also be delivered via HTTP headers, specifically the X-Robots-Tag header, configured at the server level.\nExample (Apache .htaccess):\n\u0026lt;IfModule mod_headers.c\u0026gt; Header set X-Robots-Tag \u0026#34;noai, noimageai\u0026#34; \u0026lt;/IfModule\u0026gt; Example (Nginx configuration):\nadd_header X-Robots-Tag \u0026#34;noai, noimageai\u0026#34;; Advantages: This method can apply directives to non-HTML content (e.g., PDFs, images, text files served directly) where placing HTML meta tags is not possible. Headers can also be set dynamically by the web application based on specific conditions. They are generally considered more robust if an AI scraper ignores robots.txt but still parses HTTP headers for such directives.\n5.3. Server-Side Blocking For more forceful prevention, server configurations can be used to block requests based on user-agent strings or IP addresses.\nUser-Agent Blocking: Web servers like Nginx or Apache can be configured to identify requests from specific AI bot user-agent strings and deny them access, typically by returning an HTTP 403 Forbidden status code or a 444 Connection Closed Without Response (Nginx specific).\nNginx example:\nmap $http_user_agent $block_ai_bot { default 0; ~*GPTBot 1; ~*ClaudeBot 1; # Add other AI bot UAs } server { if ($block_ai_bot) { return 403; } #... other server config... } IP Address Blocking: If known IP address ranges for AI crawlers are available (some companies like OpenAI and Perplexity publish them 6), these can be blocked at the server firewall or web server level. Considerations: These methods are more complex to implement and maintain. IP addresses can change, requiring constant updates to blocklists. User-agent strings can be spoofed, potentially leading to legitimate users being blocked if rules are not carefully crafted. This approach moves beyond polite requests into active prevention.\n5.4. Web Application Firewalls (WAFs) and Content Delivery Networks (CDNs) Commercial WAFs and CDNs (e.g., Cloudflare, AWS WAF, Akamai Bot Manager) often provide advanced bot management capabilities.\nThese systems can identify and block unwanted bots based on a variety of signals, including IP reputation, known user-agent strings, behavioral analysis (how a client interacts with a site), and machine learning models to detect sophisticated bot activity. Some CDNs offer specific features tailored to block AI scrapers. For example, Cloudflare provides an \u0026ldquo;AI Scrapers and Crawlers\u0026rdquo; blocking feature as part of its bot management solutions. Considerations: WAF/CDN solutions are typically paid services and represent a more sophisticated, often automated, layer of defense. They can be highly effective against bots that ignore robots.txt or employ evasive techniques. The array of control methods, from the simple robots.txt file to sophisticated WAFs, effectively forms an escalation path for webmasters. Typically, the simplest and most standardized methods like robots.txt are implemented first. If these prove insufficient—for instance, if a particular AI crawler ignores robots.txt and causes excessive server load (as sometimes reported for Bytespider) or if there are persistent concerns about specific content being used for training despite robots.txt directives 13—then a webmaster might progress to implementing HTML meta tags or HTTP headers. Continued non-compliance or more aggressive crawling might then warrant server-side blocking or investment in a WAF. This progression reflects a cost-benefit analysis where the perceived \u0026ldquo;cost\u0026rdquo; of AI crawling (in terms of server resources, potential content misuse, or intellectual property concerns) is weighed against the \u0026ldquo;cost\u0026rdquo; of implementing more complex controls (in terms of time, technical expertise, or financial outlay for commercial solutions). The AI industry\u0026rsquo;s overall level of respect for foundational protocols like robots.txt directly influences how quickly and how far webmasters need to escalate their defense mechanisms.\n5.5. Table: Comparison of AI Crawler Control Mechanisms The following table provides a comparative overview of the different methods discussed.\nTable 2: Comparison of AI Crawler Control Mechanisms\nMethod Implementation Level Granularity Enforcement Primary Mechanism Key Pros Key Cons robots.txt Site-wide (root) Path-based Cooperative (Polite) User-agent, Disallow directives Standardized, easy to implement, widely understood by compliant bots. Relies on bot compliance, not for security, public, can be ignored by malicious or poorly configured bots. HTML Meta Tags Page (\u0026lt;head\u0026gt;) Page-level Cooperative \u0026lt;meta name=\u0026quot;robots\u0026quot; content=\u0026quot;noai, noimageai, noml\u0026quot;\u0026gt; Page-specific control, easy for content editors. Not yet standardized, limited adoption/respect by AI bots, only for HTML documents. HTTP X-Robots-Tag Header Server (per request) Page-level Cooperative X-Robots-Tag: noai, noimageai Page-specific, works for non-HTML files, can be set dynamically by server. Not yet standardized for AI directives, relies on bot parsing headers for these specific tags. Server-Side UA Blocking Server config Site/Path Forceful Nginx/Apache rules to block UA strings, return 403/444 More effective against non-compliant bots for specific known User-Agents. Complex to maintain, risk of blocking legitimate users if UAs are spoofed, requires server configuration access. Server-Side IP Blocking Server/Firewall IP-based Forceful Firewall rules, .htaccess deny IP Effective against known bad IPs/ranges. IP addresses can change, requiring updated lists; can inadvertently block legitimate users on shared IPs. WAF/CDN Bot Management Network Edge/Server Various Forceful/Cooperative Signature, behavior, ML-based detection \u0026amp; blocking Advanced detection, can stop sophisticated/non-compliant bots, often automated. Typically paid services, configuration can be complex, potential for false positives if not tuned correctly. 6. Limitations and Best Practices While the tools discussed offer varying degrees of control, it is crucial to understand their limitations and adhere to best practices for effective AI crawler management.\n6.1. robots.txt is a Directive, Not an Enforcement Mechanism It must be reiterated that robots.txt functions based on the voluntary cooperation of web crawlers. Malicious bots, or even poorly programmed legitimate bots, can and do ignore its directives. Therefore, robots.txt should never be used as a sole method to protect sensitive or private information from being accessed. For true security, measures like password protection, server-level authentication, or IP access control lists are necessary.\n6.2. Importance of Regular Review and Updates The landscape of AI crawlers, including their user-agent tokens and crawling behaviors, is dynamic and constantly evolving. New bots emerge, and existing ones may change their identifiers or purposes. Consequently, the robots.txt file, along with any other control mechanisms, should be regularly reviewed and updated. Subscribing to industry newsletters, monitoring webmaster forums, and utilizing services that track bot activity (e.g., Dark Visitors) can help webmasters stay informed about new AI crawlers that may need to be added to their blocking rules. This transforms robots.txt management from a one-time setup into an ongoing operational task, akin to software patching or security monitoring, for those serious about comprehensive AI crawler control.\n6.3. Testing robots.txt Changes Before deploying any changes to a live robots.txt file, thorough testing is imperative to ensure the rules function as intended and do not inadvertently block desired crawlers, such as Googlebot or Bingbot, from important sections of the site. Tools like Google Search Console’s robots.txt Tester are invaluable for this purpose, allowing simulation of how Google\u0026rsquo;s crawlers interpret the file.\n6.4. Avoiding Common Pitfalls Syntax Errors: Typos in user-agent names, directive keywords (e.g., Disallow vs. Dissalow), or file paths can render rules ineffective or cause unintended behavior. Paths in robots.txt are generally case-sensitive, and user-agent tokens may also be, depending on the crawler\u0026rsquo;s implementation. Over-blocking: Care must be taken not to accidentally block search engine crawlers from content that should be indexed. The query specifically requires that search engines not be prevented from accessing the pages AI crawlers are blocked from. Blocking Essential Resources (CSS/JS): While less directly relevant to blocking AI crawlers from specific data paths, a general best practice is to avoid blocking CSS or JavaScript files that are necessary for search engines to correctly render and understand page content. Blocking these can negatively impact how search engines perceive and rank pages. Misuse of User-agent: * with Disallow: /: Applying Disallow: / to User-agent: * will block all compliant crawlers, including all search engines, from the entire site. This is directly contrary to the goal of allowing search engine access and should be avoided unless that is the specific, fully understood intention. 6.5. Log File Analysis Regular analysis of server log files is a crucial practice. Logs provide empirical data on which bots are actually crawling the site, what resources they are accessing, their request frequency, and whether they appear to be respecting robots.txt directives. This analysis can help identify:\nUnknown or new AI crawlers whose user-agent strings are not yet in the robots.txt file. Bots that are ignoring robots.txt directives, which may necessitate escalating to server-side blocking or WAF rules. Excessive crawling activity from specific bots that might be straining server resources. The \u0026ldquo;politeness\u0026rdquo; inherent in the robots.txt protocol can, unfortunately, be exploited. A sophisticated AI data scraper, aiming to circumvent specific blocks targeting known AI user-agents, could deliberately employ a generic, non-descript user-agent string (e.g., a common browser user-agent) or rotate through a list of such strings. By doing so, it would not match specific AI bot rules (like User-agent: GPTBot Disallow: /sensitive-data/) and would instead fall under the purview of any User-agent: * rules. Since webmasters are often cautious about making User-agent: * rules too restrictive to ensure broad search engine compatibility (e.g., User-agent: * Disallow: /cgi-bin/ might be common, but User-agent: * Disallow: /sensitive-data/ would block search engines too), such an evasive scraper could gain access. This highlights a fundamental vulnerability of relying solely on user-agent-based blocking within robots.txt against determined or deceptive actors and underscores the value of behavioral analysis tools or WAFs for more robust defense against such tactics.\n7. Conclusion: Implementing a Robust, Layered AI Crawler Defense Effectively managing AI crawler access while preserving search engine visibility requires a multi-layered approach, with robots.txt serving as the foundational component for communicating crawling preferences to compliant bots. This protocol, through carefully crafted User-agent and Disallow directives, allows webmasters to instruct known AI crawlers to avoid specific pages or directories.\nHowever, the reliance of robots.txt on voluntary compliance and the challenge of identifying all current and future AI crawlers mean that it is not a foolproof solution. For more comprehensive control, particularly against non-compliant bots or for highly sensitive content, webmasters should consider augmenting robots.txt with additional measures. These can include page-level HTML meta tags (such as noai or noml) and corresponding HTTP X-Robots-Tag headers as emerging standards for signaling data usage preferences for AI. For more assertive blocking, server-side configurations targeting user-agent strings or IP addresses, and sophisticated Web Application Firewalls or CDN-based bot management solutions, offer stronger enforcement capabilities.\nThe AI crawler landscape is dynamic. New bots are continuously developed, and existing ones may alter their behavior or identifiers. Therefore, ongoing vigilance, regular review of robots.txt files and server logs, and adaptation of control strategies are essential for maintaining the desired level of governance over how automated agents interact with web content.\nThe effort to block \u0026ldquo;all possible AI\u0026rdquo; while ensuring full access for search engines underscores a growing tension in web standards. The robots.txt protocol, conceived in a simpler era of web crawling, is being tested by the diverse intentions and capabilities of modern bots. This is driving the web community towards developing more nuanced signaling mechanisms for data usage (like the proposed noai tags) and compelling the adoption of more robust enforcement tools when polite directives are insufficient.\nBelow is an example robots.txt configuration designed to prevent a comprehensive list of known AI crawlers from accessing specified sections of a site, while ensuring that standard search engine bots are not similarly restricted from those sections.\n# robots.txt: Preventing AI Crawlers from Specific Content # Last Updated: October 26, 2023 - Regular review and updates are highly recommended. # ---------------------------------------------------------------------- # AI CRAWLER BLOCKING FOR SPECIFIC SECTIONS # # The following rules block specific AI crawlers from accessing: # - The entire /private-content/ directory # - The entire /research-data/ directory # - The specific file /documents/sensitive-document.pdf # ---------------------------------------------------------------------- User-agent: GPTBot Disallow: /private-content/ Disallow: /research-data/ Disallow: /documents/sensitive-document.pdf User-agent: ChatGPT-User Disallow: /private-content/ Disallow: /research-data/ Disallow: /documents/sensitive-document.pdf User-agent: OAI-SearchBot Disallow: /private-content/ Disallow: /research-data/ Disallow: /documents/sensitive-document.pdf User-agent: Google-Extended Disallow: /private-content/ Disallow: /research-data/ Disallow: /documents/sensitive-document.pdf User-agent: ClaudeBot Disallow: /private-content/ Disallow: /research-data/ Disallow: /documents/sensitive-document.pdf User-agent: anthropic-ai Disallow: /private-content/ Disallow: /research-data/ Disallow: /documents/sensitive-document.pdf User-agent: CCBot Disallow: /private-content/ Disallow: /research-data/ Disallow: /documents/sensitive-document.pdf # PerplexityBot is for indexing for Perplexity Search, not for AI model training. # Perplexity-User is for live retrieval during user queries and ignores robots.txt. # Blocking PerplexityBot from these specific sections is included here as a comprehensive measure # if any form of indexing by them on these paths is undesired. User-agent: PerplexityBot Disallow: /private-content/ Disallow: /research-data/ Disallow: /documents/sensitive-document.pdf User-agent: Bytespider # Note: Bytespider has been reported to sometimes ignore robots.txt. Disallow: /private-content/ Disallow: /research-data/ Disallow: /documents/sensitive-document.pdf User-agent: Meta-ExternalAgent Disallow: /private-content/ Disallow: /research-data/ Disallow: /documents/sensitive-document.pdf User-agent: Applebot-Extended Disallow: /private-content/ Disallow: /research-data/ Disallow: /documents/sensitive-document.pdf User-agent: cohere-ai Disallow: /private-content/ Disallow: /research-data/ Disallow: /documents/sensitive-document.pdf User-agent: GoogleOther # Google\u0026#39;s user agent for various purposes, may include R\u0026amp;D/training. Disallow: /private-content/ Disallow: /research-data/ Disallow: /documents/sensitive-document.pdf # Add other AI crawlers as they are identified, following the same pattern. # Example for a hypothetical new AI bot: # User-agent: FutureAICrawler # Disallow: /private-content/ # Disallow: /research-data/ # Disallow: /documents/sensitive-document.pdf # ---------------------------------------------------------------------- # SEARCH ENGINE CRAWLER ACCESS # # Standard search engine crawlers (Googlebot, Bingbot, DuckDuckBot, etc.) # are NOT blocked from /private-content/, /research-data/, or # /documents/sensitive-document.pdf by the rules above. # This is because their user-agent strings do not match the AI-specific # user-agents listed in the Disallow blocks. # # By default (implicit allowance), if no specific Disallow rule targets # a search engine bot for these paths, it is allowed to crawl them. # # No explicit \u0026#39;Allow:\u0026#39; rules are needed for these paths for search engines # in this specific scenario, as we are only adding Disallow rules for AI bots. # ---------------------------------------------------------------------- # Example: General rules applicable to ALL crawlers (User-agent: *) # Use with caution. These rules apply to search engines as well. # User-agent: * # Disallow: /admin/ # Example: Disallow access to an admin section for all bots. # Disallow: /tmp/ # Example: Disallow access to a temporary files folder. # Disallow: /*?sessionid= # Example: Disallow URLs with session IDs. # ---------------------------------------------------------------------- # SITEMAP DECLARATION # It is a best practice to declare the location of your XML sitemap(s). # ---------------------------------------------------------------------- Sitemap: https://www.example.com/sitemap.xml # If you use a sitemap index file, point to that: # Sitemap: https://www.example.com/sitemap_index.xml # End of robots.txt ","permalink":"https://meshrefine.com/en/conversations/robots_txt/","summary":"\u003cul\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#1-introduction-the-imperative-of-ai-crawler-management\"\u003e1. Introduction: The Imperative of AI Crawler Management\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#2-understanding-robotstxt-the-foundation-of-crawler-instruction\"\u003e2. Understanding \u003ccode\u003erobots.txt\u003c/code\u003e: The Foundation of Crawler Instruction\u003c/a\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#21-core-syntax-and-directives\"\u003e2.1. Core Syntax and Directives\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#22-file-placement-and-formatting\"\u003e2.2. File Placement and Formatting\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#23-how-crawlers-interpret-robotstxt\"\u003e2.3. How Crawlers Interpret \u003ccode\u003erobots.txt\u003c/code\u003e\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#24-testing-your-robotstxt\"\u003e2.4. Testing Your \u003ccode\u003erobots.txt\u003c/code\u003e\u003c/a\u003e\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#3-identifying-key-crawler-types-ai-agents-vs-search-engine-bots\"\u003e3. Identifying Key Crawler Types: AI Agents vs. Search Engine Bots\u003c/a\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#31-distinguishing-characteristics\"\u003e3.1. Distinguishing Characteristics\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#32-categories-of-ai-crawlers-and-their-user-agents\"\u003e3.2. Categories of AI Crawlers and Their User Agents\u003c/a\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#321-ai-crawlers-for-model-training\"\u003e3.2.1. AI Crawlers for Model Training\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#322-ai-crawlers-for-live-retrieval-and-search-assistance\"\u003e3.2.2. AI Crawlers for Live Retrieval and Search Assistance\u003c/a\u003e\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#33-standard-search-engine-crawlers-to-be-allowed\"\u003e3.3. Standard Search Engine Crawlers (to be Allowed)\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#34-table-of-prominent-ai-crawler-user-agents\"\u003e3.4. Table of Prominent AI Crawler User Agents\u003c/a\u003e\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#4-strategically-blocking-ai-crawlers-with-robotstxt\"\u003e4. Strategically Blocking AI Crawlers with \u003ccode\u003erobots.txt\u003c/code\u003e\u003c/a\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#41-targeting-specific-ai-user-agents\"\u003e4.1. Targeting Specific AI User Agents\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#42-applying-rules-to-specific-pages-or-directories\"\u003e4.2. Applying Rules to Specific Pages or Directories\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#43-ensuring-search-engines-are-not-blocked-from-specific-pages\"\u003e4.3. Ensuring Search Engines Are Not Blocked from Specific Pages\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#44-the-challenge-of-all-possible-ai\"\u003e4.4. The Challenge of \u0026ldquo;All Possible AI\u0026rdquo;\u003c/a\u003e\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#5-advanced-methods-for-granular-ai-crawler-control\"\u003e5. Advanced Methods for Granular AI Crawler Control\u003c/a\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#51-html-meta-tags-page-level-control\"\u003e5.1. HTML Meta Tags (Page-Level Control)\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#52-http-x-robots-tag-headers-server-level-page-control\"\u003e5.2. HTTP \u003ccode\u003eX-Robots-Tag\u003c/code\u003e Headers (Server-Level Page Control)\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#53-server-side-blocking\"\u003e5.3. Server-Side Blocking\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#54-web-application-firewalls-wafs-and-content-delivery-networks-cdns\"\u003e5.4. Web Application Firewalls (WAFs) and Content Delivery Networks (CDNs)\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#55-table-comparison-of-ai-crawler-control-mechanisms\"\u003e5.5. Table: Comparison of AI Crawler Control Mechanisms\u003c/a\u003e\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#6-limitations-and-best-practices\"\u003e6. Limitations and Best Practices\u003c/a\u003e\n\u003cul\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#61-robotstxt-is-a-directive-not-an-enforcement-mechanism\"\u003e6.1. \u003ccode\u003erobots.txt\u003c/code\u003e is a Directive, Not an Enforcement Mechanism\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#62-importance-of-regular-review-and-updates\"\u003e6.2. Importance of Regular Review and Updates\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#63-testing-robotstxt-changes\"\u003e6.3. Testing \u003ccode\u003erobots.txt\u003c/code\u003e Changes\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#64-avoiding-common-pitfalls\"\u003e6.4. Avoiding Common Pitfalls\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#65-log-file-analysis\"\u003e6.5. Log File Analysis\u003c/a\u003e\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\u003ca href=\"/en/conversations/robots_txt/#7-conclusion-implementing-a-robust-layered-ai-crawler-defense\"\u003e7. Conclusion: Implementing a Robust, Layered AI Crawler Defense\u003c/a\u003e\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch2 id=\"1-introduction-the-imperative-of-ai-crawler-management\"\u003e1. Introduction: The Imperative of AI Crawler Management\u003c/h2\u003e\n\u003cp\u003eThe proliferation of Artificial Intelligence (AI) has introduced a new class of web crawlers designed to gather vast quantities of data for training Large Language Models (LLMs) and powering AI-driven applications. While these advancements offer significant potential, website operators often require precise control over which content AI crawlers can access, particularly to protect intellectual property, sensitive information, or manage server resources. Simultaneously, maintaining visibility and crawlability for traditional search engine bots like Googlebot and Bingbot remains paramount for organic search performance.\u003c/p\u003e","title":"Mastering AI Crawler Control: A Guide to `robots.txt` and Advanced Webmaster Tools"},{"content":" My name is Konstantin Meshcheryakov. I\u0026rsquo;m Head of AI at an IoT engineering company in Wroclaw, where I build the GenAI strategy, run the AI adoption program for the whole engineering organisation, and own the company\u0026rsquo;s EU AI Act compliance work. I also keep doing architecture on AI engagements, since advice ages quickly without delivery behind it.\nThe road here took sixteen years. I started with a PhD in Physics and Mathematics, wrote C++/Qt systems for aircraft testing, moved through embedded development and AWS architecture, and picked up machine learning along the way, from TinyML on microcontrollers to production ML and LLM systems. For several years I also lectured at a university, teaching Qt, Python, and functional programming. That arc, from firmware to GenAI strategy, still shapes how I judge what will survive contact with production. The blog\u0026rsquo;s name is borrowed from the CFD vocabulary, with a convenient overlap with my surname.\nThis blog is where I share my thoughts and observations from the AI trenches: practical use cases, nuances of working with LLMs and agents, AI-assisted development, and the pitfalls of edge AI. Occasionally, you might also find notes here on other tech topics that catch my attention. All of this reflects my personal experience and perspective, seasoned with a touch of irony and a commitment to separating substance from hype.\n","permalink":"https://meshrefine.com/en/about/","summary":"\u003cimg src=\"/images/konstantin-meshcheryakov.jpg\" alt=\"Konstantin Meshcheryakov\" style=\"float: right; width: 200px; max-width: 45%; height: auto; border-radius: 12px; margin: 0.25rem 0 1rem 1.5rem;\"\u003e\n\u003cp\u003eMy name is Konstantin Meshcheryakov. I\u0026rsquo;m Head of AI at an IoT engineering company in Wroclaw, where I build the GenAI strategy, run the AI adoption program for the whole engineering organisation, and own the company\u0026rsquo;s EU AI Act compliance work. I also keep doing architecture on AI engagements, since advice ages quickly without delivery behind it.\u003c/p\u003e","title":"About Me"},{"content":"On this website, I am committed to full transparency regarding the use of Artificial Intelligence (AI) technologies. All content is clearly divided into the following categories:\nAuthor-Written Content. The majority of the content on this site is written by me personally. AI may be used solely for basic assistance, such as spell-checking and grammar correction.\nLabeling: This content is not specifically labeled. AI Access: Permitted for search engine indexing and for use in training datasets for AI models. AI-Translated Content. Articles that are translations of my original content into another language. The initial translation is performed by AI, after which I personally proofread and edit the text to ensure accuracy and stylistic consistency.\nLabeling: Such materials are explicitly marked to indicate the use of AI for translation (e.g., \u0026ldquo;Translated with AI\u0026rdquo;). AI Access: Permitted for indexing and AI training. I believe that the human editing and the original content base provide sufficient quality to avoid model degradation. AI-Generated Content. Materials where AI was used to generate the main body of the text, which was then thoroughly reviewed, corrected, and supplemented by me.\nLabeling: This content is explicitly marked (e.g., \u0026ldquo;Generated with AI\u0026rdquo;). AI Access: Permitted for search engine indexing but is prohibited from use in training AI models (noai). ","permalink":"https://meshrefine.com/en/ai-content-policy/","summary":"\u003cp\u003eOn this website, I am committed to full transparency regarding the use of Artificial Intelligence (AI) technologies. All content is clearly divided into the following categories:\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003e\n\u003cp\u003eAuthor-Written Content. The majority of the content on this site is written by me personally. AI may be used solely for basic assistance, such as spell-checking and grammar correction.\u003c/p\u003e\n\u003cul\u003e\n\u003cli\u003eLabeling: This content is not specifically labeled.\u003c/li\u003e\n\u003cli\u003eAI Access: Permitted for search engine indexing and for use in training datasets for AI models.\u003c/li\u003e\n\u003c/ul\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAI-Translated Content. Articles that are translations of my original content into another language. The initial translation is performed by AI, after which I personally proofread and edit the text to ensure accuracy and stylistic consistency.\u003c/p\u003e","title":"AI Content Policy"},{"content":"Today we will look at Griptape, a framework for building AI applications, which offers a clean Pythonic API for those tired of LangChain\u0026rsquo;s abstraction layers. It provides primitives for building assistants, RAG systems, and integrating with external tools. Honestly, in my experience, most people tired of LangChain switch to custom-written wrappers around lower-level libraries like OpenAI or LiteLLM. But who knows, maybe they\u0026rsquo;re missing out. Let\u0026rsquo;s dive in.\nA Bit of History Personally, I\u0026rsquo;ve been hearing about Griptape for about a year and a half. As far as I remember, It started as a sort of LangChain competitor with quite similar primitives, but their paths gradually diverged. As of the time of the writing, it has 2.3k stars on GitHub, which is somewhat less than LangChain\u0026rsquo;s 109k, but still enough to consider the project quite mature. Besides the open-source framework, it has also developed its own cloud where you can run your applications, ETLs, and RAGs, and a visual builder, Griptape Nodes, allowing non-professionals to click together applications in minutes. 1\nThe Framework Let\u0026rsquo;s get to the framework itself. It introduces several primitives that are worth understanding before starting the actual use.\nDrivers: Essentially the main tool, abstracting specific implementations of anything and allowing them to be changed on the fly without breaking business logic. There are drivers for almost everything, be it assistants, prompts, models, embedding systems, or databases. Basically, they look (and implemented) as abstract classes and their implementations.\nEngines: Provide ready-made implementations for core tasks such as RAG and summarization.\nStructures: Basic building blocks for constructing applications. Among them, you can find:\nTasks: An abstraction over some action, for example, querying a model, processing a response, loading data from a file, and so on.\nAgents: A somewhat strange wrapper around a single task 2. It allows passing input and a list of tools that this task can use.\nPipelines: Like an agent, but can run multiple tasks sequentially, passing the output of one to the input of another.\nWorkflows: Directed acyclic graphs (DAGs) consisting of tasks. Allows for optimal planning and execution of parallel tasks. The documentation states they are non-sequential, although it later provides examples of quite sequential workflows. In this case, it\u0026rsquo;s somewhat unclear why pipelines are needed at all.\nOverall, this organization is quite flexible and allows for creating fairly complex execution flows. The fact that it\u0026rsquo;s a DAG imposes certain limitations on creating agents in the sense that some futuristically-minded people envision, but it allows for building reliable 3 and well-thought-out systems.\nTools: Functions available to LLMs. These functions can be passed to the primitives mentioned above, thus giving the LLM the ability to generate a sequence of calls to these functions to perform actions. Griptape provides quite a few such tools, but you can also add your own.\nTask Memory: One of Griptape\u0026rsquo;s interesting features. Often, the data to be processed is either very sensitive or very large, and sending it directly to the LLM can be impractical. In such cases, you can ask the tool not to provide this data to the LLM, but to return some descriptor of this data that allows the LLM to reference it for use in other tools.\nConversation Memory: Enabled by default, and passed to the model between runs of the same workflow.4 Can be disabled if not required. Essentially, useful only for chatbots and potentially harmful for everything else.\nRulesets: LLM behavior settings that are passed into every prompt. Something like system prompts, but to understand exactly what they are, one needs to dig deeper.\nThat covers the main primitives; there are a few other concepts that aren\u0026rsquo;t too important at this stage, so I\u0026rsquo;ll skip them.\nWhat\u0026rsquo;s Next? Next, I plan to dig deeper into:\nWhen and why to use agents, pipelines, and workflows The framework\u0026rsquo;s capabilities for ETL and RAG How Off-Prompt Memory works Integration and customization What can be done in their cloud And also what Griptape Nodes is good for And, of course, run some of the examples they provide. I\u0026rsquo;ll describe my progress as I go.\nnot really\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nalthough there is an API for creating an agent around a list of tasks, it fails when trying to pass more than one task, which is quite strange.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nas reliable as possible with probabilistic LLMs\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nor agent, or pipeline\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://meshrefine.com/en/posts/griptape-1/","summary":"\u003cp\u003eToday we will look at \u003ca href=\"www.griptape.ai\"\u003eGriptape\u003c/a\u003e, a framework for building AI applications, which offers a clean Pythonic API for those tired of LangChain\u0026rsquo;s abstraction layers. It provides primitives for building assistants, RAG systems, and integrating with external tools. Honestly, in my experience, most people tired of LangChain switch to custom-written wrappers around lower-level libraries like OpenAI or LiteLLM. But who knows, maybe they\u0026rsquo;re missing out. Let\u0026rsquo;s dive in.\u003c/p\u003e\n\u003ch1 id=\"a-bit-of-history\"\u003eA Bit of History\u003c/h1\u003e\n\u003cp\u003ePersonally, I\u0026rsquo;ve been hearing about Griptape for about a year and a half. As far as I remember, It started as a sort of LangChain competitor with quite similar primitives, but their paths gradually diverged. As of the time of the writing, it has 2.3k stars on GitHub, which is somewhat less than LangChain\u0026rsquo;s 109k, but still enough to consider the project quite mature.\nBesides the open-source framework, it has also developed its own cloud where you can run your applications, ETLs, and RAGs, and a visual builder, Griptape Nodes, allowing non-professionals to click together applications in minutes. \u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e\u003c/p\u003e","title":"Griptape: A Framework for AI Applications, Part 1: Introduction"},{"content":"Alright, Here We Go I got my hands on the Re:Camera from Seeed. Essentially, it\u0026rsquo;s a small box (a cube about 4 cm per side), wrapped in a heatsink. Inside, there\u0026rsquo;s a dual-core RISC-V based MPU (updated: only one core is visible to the system, the second one is apparently reserved for special operations), an ancient 8051 microcontroller, an OmniVision camera sensor, LEDs for illumination, Wi-Fi, BT, and, you know, all sorts of peripherals. RAM is a bit scarce, only 256 megabytes, so getting Greengrass on it will be problematic. You can connect Ethernet via a special dongle-adapter that barely stays put, but for development, there\u0026rsquo;s no point, because the camera shares its network over USB type C, and it\u0026rsquo;s easier to work that way. If you\u0026rsquo;re short on storage (and the device comes in 8 GB and 64 GB built-in storage options), you can stick in a MicroSD card. You can also stick the box to something metallic, as it has magnets on one side.\nWhen connected, you can open a web page (192.168.42.1 by default), which, initially, don\u0026rsquo;t let you do anything except for providing access to the network setup and updating the system. There\u0026rsquo;s also a console there, but who in their right mind uses a console in a browser when you have SSH?\nI suppose you can do something with it in its current state, but the manual insistently suggest updating via OTA, which I did (as obvious as it sounds, you need to set up the network first for that). After 5 minutes of active LED blinking and the device connecting/disconnecting from the computer, it granted me access to an updated page and required changing the default password (recamera/recamera). Changing this password also changes the SSH access password, so be careful.\nAfter the update, you get access to Node-RED, a sort of graphical development environment where data flows are described by graphs with processing nodes. By default, such a graph is already loaded and gives access to a simple dashboard that lets you see the camera image and the results of the YOLOv11 AI model. Other graphs can (in theory) be grabbed from Seeed\u0026rsquo;s SenseCraft AI website, but it requires registration, so I couldn\u0026rsquo;t be bothered.\nAbout the Camera The sensor installed in this box is frankly outdated. It\u0026rsquo;s an OmniVision OV5647, a 5-megapixel camera with a rolling shutter. Just recently, I was working on a demo with an STMicro BrightSense sensor, equipped with a global shutter, which captures information from the entire matrix at once, not line by line like this one. This affected the resolution (only 1.5 megapixels), but on the other hand, it captured fast-falling candies perfectly. You can read about all the \u0026ldquo;joys\u0026rdquo; of a rolling shutter on Wikipedia, but in a nutshell, if you\u0026rsquo;ve seen how they beautifully depict things getting sucked into a black hole in movies, well, that\u0026rsquo;s it. Seeed promises that other sensors can be connected and that they\u0026rsquo;ll release new versions in the future, including one with a global shutter camera, but for now, it is what it is.\nThe window should be straight Why is it important to have undistorted objects? Because of the model. Which we\u0026rsquo;ll get to now.\nAbout the Model Well, not quite. First, about what it runs on. The little box has an NPU rated at a whole TOPS. When using 8-bit quantization, of course, it can\u0026rsquo;t do floats. I haven\u0026rsquo;t yet figured out how to optimize anything for this NPU, so for now, performance can only be judged by the data available on the built-in dashboard.\nThe model used there, Ultralytics YOLOv11 in its smallest variant (n, i.e., nano), can detect and output bounding boxes for 80 different classes of objects, like giraffes and toothbrushes, which doesn\u0026rsquo;t really fit most real-world tasks, so the model needs to be trained on your own dataset. But the performance is quite decent; the demo provides the following data:\nPre-processing: 0 ms. Now that\u0026rsquo;s great; apparently, the camera can output results directly in a format the model can take. In the demo I mentioned, such a setup didn\u0026rsquo;t work, and I had to tinker quite a bit to optimize it. Inference itself: ~50 ms. Not bad, not bad at all. Post-processing: 20–25 ms. This is where the Non-Maximum Suppression algorithm runs, which is quite resource-hungry, and it seems to be running on the CPU, judging by the load in top. So, all in all, this gives us something around 12–13 inferences per second, which is quite tolerable for many tasks. In the UI, however, it visually looks like 4–5 frames per second, which might be due to a non-optimal pipeline or other overhead.\nWith all this going on, the little box gets pretty warm.\nWhat\u0026rsquo;s (Still) Off-Screen If we\u0026rsquo;re talking about using this for something more serious than playing with built-in demos, we need to figure out:\na) How to build the system. I doubt Node-RED is suitable for production tasks, so you\u0026rsquo;ll need to roll up your sleeves (Buildroot, by the way) and shoehorn in the software we need.\nb) How to optimize the model and run inference from our own programs.\nIf I have time for this, I\u0026rsquo;ll come back to it in future installments. But for now, catch you later.\n","permalink":"https://meshrefine.com/en/posts/re-camera-1/","summary":"\u003ch2 id=\"alright-here-we-go\"\u003eAlright, Here We Go\u003c/h2\u003e\n\u003cp\u003eI got my hands on the Re:Camera from Seeed. Essentially, it\u0026rsquo;s a small box (a cube about 4 cm per side), wrapped in a heatsink. Inside, there\u0026rsquo;s a dual-core RISC-V based MPU (\u003cstrong\u003eupdated:\u003c/strong\u003e only one core is visible to the system, the second one is apparently reserved for special operations), an ancient 8051 microcontroller, an OmniVision camera sensor, LEDs for illumination, Wi-Fi, BT, and, you know, all sorts of peripherals. RAM is a bit scarce, only 256 megabytes, so getting Greengrass on it will be problematic. You can connect Ethernet via a special dongle-adapter that barely stays put, but for development, there\u0026rsquo;s no point, because the camera shares its network over USB type C, and it\u0026rsquo;s easier to work that way. If you\u0026rsquo;re short on storage (and the device comes in 8 GB and 64 GB built-in storage options), you can stick in a MicroSD card. You can also stick the box to something metallic, as it has magnets on one side.\u003c/p\u003e","title":"Seeed Re:Camera review, part 1"}]