{"articles":[{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"Milestone Reached: YINI Syntax Highlighting Is Now on the VS Code Marketplace","link":"https://dev.to/marko_kseppnen_6250a7f/milestone-reached-yini-syntax-highlighting-is-now-on-the-vs-code-marketplace-18mb","description":"<h2>\n\n\nMilestone Reached: YINI Syntax Highlighting Is Now on the VS Code Marketplace\n</h2>\n\n<p>A small but important YINI milestone has been reached: the official <strong>YINI Syntax Highlighting</strong> extension for Visual Studio Code has reached version <strong>1.0.0</strong> and is now available on the <strong>Visual Studio Marketplace</strong>.</p>\n\n<p>The extension targets <strong>YINI Specification 1.0.0 RC 6</strong> and adds syntax highlighting for <code>.yini</code> configuration files directly in VS Code.</p>\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7cqnwdjljv7fhii2u9l4.png\"><img alt=\" \" height=\"504\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7cqnwdjljv7fhii2u9l4.png\" width=\"615\" /></a></p>\n\n<p>It highlights YINI sections, keys, values, strings, numbers, comments, lists, objects, directives, and other RC 6 syntax.</p>\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feqt9jq74460uwmsw5b6d.png\"><img alt=\" \" height=\"760\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feqt9jq74460uwmsw5b6d.png\" width=\"800\" /></a></p>\n\n<p>You can install it directly from VS Code by searching for:</p>\n\n<p><strong>YINI Syntax Highlighting</strong></p>\n\n<p>Marketplace:</p>\n\n<p><a href=\"https://marketplace.visualstudio.com/items?itemName=yini-lang.yini-syntax-highlighting\" rel=\"noopener noreferrer\">https://marketplace.visualstudio.com/items?itemName=yini-lang.yini-syntax-highlighting</a><br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight plaintext\"><code>Version: 1.0.0\nPublisher: yini-lang\nPackage: yini-syntax-highlighting\n</code></pre>\n\n</div>\n\n\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fipyaxj0lnhi768r060bu.png\"><img alt=\" \" height=\"255\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fipyaxj0lnhi768r060bu.png\" width=\"466\" /></a></p>\n\n<p>Syntax highlighting is only one part of the YINI ecosystem, but it makes <code>.yini</code> files much nicer to read and work with in everyday development.</p>\n\n<p>More about YINI:<br />\n<strong>Homepage:</strong> <a href=\"https://yini-lang.org\" rel=\"noopener noreferrer\">https://yini-lang.org</a><br />\n<strong>GitHub:</strong> <a href=\"https://github.com/YINI-lang\" rel=\"noopener noreferrer\">https://github.com/YINI-lang</a></p>","content":"Milestone Reached: YINI Syntax Highlighting Is Now on the VS Code Marketplace\nA small but important YINI milestone has been reached: the official YINI Syntax Highlighting extension for Visual Studio Code has reached version 1.0.0 and is now available on the Visual Studio Marketplace.\nThe extension targets YINI Specification 1.0.0 RC 6 and adds syntax highlighting for .yini\nconfiguration files directly in VS Code.\nIt highlights YINI sections, keys, values, strings, numbers, comments, lists, objects, directives, and other RC 6 syntax.\nYou can install it directly from VS Code by searching for:\nYINI Syntax Highlighting\nMarketplace:\nhttps://marketplace.visualstudio.com/items?itemName=yini-lang.yini-syntax-highlighting\nVersion: 1.0.0\nPublisher: yini-lang\nPackage: yini-syntax-highlighting\nSyntax highlighting is only one part of the YINI ecosystem, but it makes .yini\nfiles much nicer to read and work with in everyday development.\nMore about YINI:\nHomepage: https://yini-lang.org\nGitHub: https://github.com/YINI-lang\nTop comments (0)","published":"Sat, 12 Sep 2026 23:31:57 +0000","author":"Mr Smithnen","guid":"https://dev.to/marko_kseppnen_6250a7f/milestone-reached-yini-syntax-highlighting-is-now-on-the-vs-code-marketplace-18mb","created_at":"2026-09-13T01:36:17.352356","last_synchronized":"2026-09-13T01:36:17.352356","sentiment":{"sentiment":"Positive","score":0.9074,"details":{"neg":0.0,"neu":0.925,"pos":0.075,"compound":0.9074}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"Building a Freelance Rate Calculator for Mexico with Plain HTML and JavaScript","link":"https://dev.to/franciscopadillamx/building-a-privacy-first-freelance-rate-calculator-with-plain-html-and-javascript-15k4","description":"<p>Freelancers in Mexico often start with a tempting formula: take the monthly income they want, divide it by the number of hours they expect to work, and publish that number as an hourly rate. The result is usually too low because it treats every working hour as billable and ignores operating costs, time off, and uncertainty.</p>\n\n<p>This article shows how I built a small, privacy-first freelance rate calculator in plain HTML and JavaScript. It makes the assumptions visible without an account, upload, framework, or server-side database.</p>\n\n<h2>\n\n\nThe model\n</h2>\n\n<p>A practical starting point is:</p>\n\n<p>billable hours = working hours × billable ratio<br />\nrequired revenue = target income + monthly expenses + contingency<br />\nhourly rate = required revenue ÷ billable hours</p>\n\n<p>The billable ratio matters. Administration, proposals, learning, breaks, and unpaid client communication all consume time. A freelancer who works 160 hours in a month may only be able to invoice 80–110 of them.</p>\n\n<p>The contingency can be a fixed amount or a percentage. A percentage is useful when expenses or income vary. For example, with a 10% buffer, calculate the base as target income plus monthly expenses, calculate 10% of that base, and add the result before dividing by billable hours.</p>\n\n<p>The implementation should handle invalid inputs before calculating. Zero or negative billable hours must not produce a misleading result, and a missing field should be visible to the user instead of silently becoming zero.</p>\n\n<h2>\n\n\nWhy plain JavaScript is a good fit\n</h2>\n\n<p>For a personal pricing calculator, a static page has useful properties:</p>\n\n<ul>\n<li>It loads quickly and can work without a framework or build step.</li>\n<li>The formula is visible and auditable in the browser.</li>\n<li>There is no account or server-side database to maintain.</li>\n<li>The inputs can stay in the browser, which is appropriate for private financial planning.</li>\n<li>The same calculation can be reused in a quote or receipt workflow.</li>\n</ul>\n\n<p>The trade-off is that a static page does not replace professional tax or accounting advice. It is a planning aid, not a promise of income or a legal calculation.</p>\n\n<h2>\n\n\nA small UX detail that improves the result\n</h2>\n\n<p>Show the intermediate values, not only the final rate. A useful result panel can display estimated billable hours, monthly revenue required, contingency amount, suggested hourly rate, and a short explanation of which assumptions drove the result.</p>\n\n<p>This gives the freelancer something they can challenge. If the rate is unexpectedly high, they can decide whether to change expenses, capacity, target income, or the billable ratio. That is more useful than hiding the assumptions behind a single number.</p>\n\n<h2>\n\n\nKeep the tool portable\n</h2>\n\n<p>If the calculator is used offline, avoid external fonts, analytics scripts, and runtime dependencies. Use semantic HTML, labels connected to inputs, keyboard-friendly controls, and a clear focus state. Test the page with JavaScript disabled so that the explanatory content still makes sense even when the calculation is unavailable.</p>\n\n<p>For a tool that eventually grows into quotes, receipts, or tracking, keep the calculation function separate from the DOM code. A pure function can receive target income, monthly expenses, working hours, billable ratio, and contingency ratio; it can validate billable hours, calculate the revenue requirement, and return the intermediate values and final rate. The page can then format the result without mixing business logic and presentation.</p>\n\n<p>The pure function is easy to test with a few known examples. Test a normal case, zero billable hours, an empty field, and a very low billable ratio. These tests catch more useful mistakes than checking only whether a button appears to work.</p>\n\n<h2>\n\n\nWhat I learned from making a small version\n</h2>\n\n<p>The most important design decision was keeping the free calculation useful on its own. A visitor should be able to understand the formula and get a result without creating an account or uploading financial data. Optional offline templates can help with the next step, but they should not be required to evaluate the basic idea.</p>\n\n<p>That separation also makes the project easier to maintain. The calculator can remain a small static page, while quote, receipt, and tracking workflows can evolve independently. Each one can be downloaded, inspected, and used without depending on a backend service.</p>\n\n<h2>\n\n\nTry the working example\n</h2>\n\n<p>I built a small free browser-based example for freelancers in Mexico. It runs in the browser, explains the assumptions, and links to optional offline quote and receipt templates for people who want a reusable workflow:</p>\n\n<p>Open Finanzas Freelance MX: <a href=\"https://finanzas-freelance-mx.jfpadilla1101.chatgpt.site/freelance-rate-calculator.html\" rel=\"noopener noreferrer\">https://finanzas-freelance-mx.jfpadilla1101.chatgpt.site/freelance-rate-calculator.html</a></p>\n\n<p>For a reusable offline workflow after testing the calculator, the Kit Freelance MX bundle includes a quote template, receipt, financial tracker, and guide: <a href=\"https://payhip.com/b/P8zJa?utm_source=dev&amp;utm_medium=article&amp;utm_campaign=freelance-rate-calculator\" rel=\"noopener noreferrer\">https://payhip.com/b/P8zJa?utm_source=dev&amp;utm_medium=article&amp;utm_campaign=freelance-rate-calculator</a></p>\n\n<p>The important part is the method: make capacity visible, price the work against the revenue requirement, and revisit the assumptions when reality changes.</p>\n\n<p>Disclosure: this article was prepared with AI assistance, and the linked tool and templates are my own project. The article is intended to be useful independently of the link.</p>","content":"Freelancers in Mexico often start with a tempting formula: take the monthly income they want, divide it by the number of hours they expect to work, and publish that number as an hourly rate. The result is usually too low because it treats every working hour as billable and ignores operating costs, time off, and uncertainty.\nThis article shows how I built a small, privacy-first freelance rate calculator in plain HTML and JavaScript. It makes the assumptions visible without an account, upload, framework, or server-side database.\nThe model\nA practical starting point is:\nbillable hours = working hours × billable ratio\nrequired revenue = target income + monthly expenses + contingency\nhourly rate = required revenue ÷ billable hours\nThe billable ratio matters. Administration, proposals, learning, breaks, and unpaid client communication all consume time. A freelancer who works 160 hours in a month may only be able to invoice 80–110 of them.\nThe contingency can be a fixed amount or a percentage. A percentage is useful when expenses or income vary. For example, with a 10% buffer, calculate the base as target income plus monthly expenses, calculate 10% of that base, and add the result before dividing by billable hours.\nThe implementation should handle invalid inputs before calculating. Zero or negative billable hours must not produce a misleading result, and a missing field should be visible to the user instead of silently becoming zero.\nWhy plain JavaScript is a good fit\nFor a personal pricing calculator, a static page has useful properties:\n- It loads quickly and can work without a framework or build step.\n- The formula is visible and auditable in the browser.\n- There is no account or server-side database to maintain.\n- The inputs can stay in the browser, which is appropriate for private financial planning.\n- The same calculation can be reused in a quote or receipt workflow.\nThe trade-off is that a static page does not replace professional tax or accounting advice. It is a planning aid, not a promise of income or a legal calculation.\nA small UX detail that improves the result\nShow the intermediate values, not only the final rate. A useful result panel can display estimated billable hours, monthly revenue required, contingency amount, suggested hourly rate, and a short explanation of which assumptions drove the result.\nThis gives the freelancer something they can challenge. If the rate is unexpectedly high, they can decide whether to change expenses, capacity, target income, or the billable ratio. That is more useful than hiding the assumptions behind a single number.\nKeep the tool portable\nIf the calculator is used offline, avoid external fonts, analytics scripts, and runtime dependencies. Use semantic HTML, labels connected to inputs, keyboard-friendly controls, and a clear focus state. Test the page with JavaScript disabled so that the explanatory content still makes sense even when the calculation is unavailable.\nFor a tool that eventually grows into quotes, receipts, or tracking, keep the calculation function separate from the DOM code. A pure function can receive target income, monthly expenses, working hours, billable ratio, and contingency ratio; it can validate billable hours, calculate the revenue requirement, and return the intermediate values and final rate. The page can then format the result without mixing business logic and presentation.\nThe pure function is easy to test with a few known examples. Test a normal case, zero billable hours, an empty field, and a very low billable ratio. These tests catch more useful mistakes than checking only whether a button appears to work.\nWhat I learned from making a small version\nThe most important design decision was keeping the free calculation useful on its own. A visitor should be able to understand the formula and get a result without creating an account or uploading financial data. Optional offline templates can help with the next step, but they should not be required to evaluate the basic idea.\nThat separation also makes the project easier to maintain. The calculator can remain a small static page, while quote, receipt, and tracking workflows can evolve independently. Each one can be downloaded, inspected, and used without depending on a backend service.\nTry the working example\nI built a small free browser-based example for freelancers in Mexico. It runs in the browser, explains the assumptions, and links to optional offline quote and receipt templates for people who want a reusable workflow:\nOpen Finanzas Freelance MX: https://finanzas-freelance-mx.jfpadilla1101.chatgpt.site/freelance-rate-calculator.html\nFor a reusable offline workflow after testing the calculator, the Kit Freelance MX bundle includes a quote template, receipt, financial tracker, and guide: https://payhip.com/b/P8zJa?utm_source=dev&utm_medium=article&utm_campaign=freelance-rate-calculator\nThe important part is the method: make capacity visible, price the work against the revenue requirement, and revisit the assumptions when reality changes.\nDisclosure: this article was prepared with AI assistance, and the linked tool and templates are my own project. The article is intended to be useful independently of the link.\nTop comments (0)","published":"Sat, 12 Sep 2026 23:12:27 +0000","author":"Francisco Padilla","guid":"https://dev.to/franciscopadillamx/building-a-privacy-first-freelance-rate-calculator-with-plain-html-and-javascript-15k4","created_at":"2026-09-13T01:36:17.352356","last_synchronized":"2026-09-13T01:36:17.352356","sentiment":{"sentiment":"Positive","score":0.9867,"details":{"neg":0.03,"neu":0.89,"pos":0.08,"compound":0.9867}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"I built a chat where every word costs money — on TON, solo, no legal entity","link":"https://dev.to/__e3096294/i-built-a-chat-where-every-word-costs-money-on-ton-solo-no-legal-entity-1f3f","description":"<p>Hi! I'm a mobile developer (Kotlin by day), and a couple of weeks ago I started a side project with a simple, slightly audacious idea: <strong>a chat where sending a message costs money — and the recipient gets paid</strong>. You write \"hi\", you pay. Someone writes to you, you earn.</p>\n\n<p>This is the story of how \"what if attention literally had a price\" became a working Telegram Mini App with real on-chain transactions — and the rakes I stepped on along the way, from app store rules to the anatomy of TON gas. An LLM wrote the code with me, we argued about architecture together, but the debugging bills were mine — literally, since debugging happened on real money.</p>\n\n<h2>\n\n\nThe idea: a market of expensive attention\n</h2>\n\n<p>Free messages are worthless — in both senses. Spam, \"hey what's up\" from strangers, group chats with 10,000 unread messages. What if we flip it: <strong>every character costs money</strong>?</p>\n\n<p>The economics get interesting fast:</p>\n\n<ul>\n<li>you can message any stranger — but you'll pay <em>them</em>. A message with money inside always gets opened;</li>\n<li>the public chat becomes an auction of wit: the priciest message of the week wears the crown 👑 and stays pinned;</li>\n<li>spam dies on its own: spamming at $0.05 per character is a fast way to go broke.</li>\n</ul>\n\n<p>I called it Centence. Between the idea and the product stood two walls.</p>\n\n<h2>\n\n\nWall one: the app stores\n</h2>\n\n<p>A mobile developer's first reflex is a native app. It's also the last one: on iOS and Android any payment for \"digital content\" must go through in-app purchases — 30% commission and rules that treat P2P money transfers between users as a minefield, and crypto as a minefield in fog.</p>\n\n<p>Option two: a Telegram Mini App. But there's a rule there too — digital goods and services inside mini apps must be sold for Telegram Stars. Paying for \"a message\" with Stars, with the platform taking a cut, killed the whole point: money must go <strong>to the recipient</strong>, directly.</p>\n\n<p>Then came the pleasant realization: the Stars rule covers <em>buying digital goods from the app</em>. If the app sells nothing and never touches money — if it's just an interface through which one person transfers money from their wallet to another person's wallet — that's not a purchase. That's a transfer.</p>\n\n<h2>\n\n\nWall two: the legal entity\n</h2>\n\n<p>The moment a service accepts users' money into its own account (even \"for a second\", even \"just to forward it\"), it becomes a money operator. That means licenses, KYC/AML, a legal entity — everything that kills a side project on day one.</p>\n\n<p>Hence the architectural principle that everything else obeys:</p>\n\n<blockquote>\n<p><strong>The platform never touches user funds.</strong> No app balance, no deposits, no withdrawals, no escrow, no refunds. The server holds no private keys and is physically incapable of moving anyone's money.</p>\n</blockquote>\n\n<p>A payment goes <strong>from the sender's wallet straight to the recipient's wallet</strong> via TON Connect. My fee (10% on direct messages) is a separate output of the same transaction, to my own wallet. If the server gets hacked — there's no money in it. None.</p>\n\n<p>Legal bonus: since we're not a money operator, no licenses and no legal entity needed. Taxes on the fee are ordinary personal income.</p>\n\n<h2>\n\n\nThe architecture: matching a transaction to a message\n</h2>\n\n<p>The central technical problem of a non-custodial design: the server must figure out that a specific on-chain transaction pays for a specific message — while the transaction is signed by the user in their own wallet, and until the network confirms it, no \"payment\" exists.</p>\n\n<p>The scheme:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight plaintext\"><code>1. Client sends the text to the server → server computes the price,\n stores {uuid, text, amounts, status: pending}\n2. Client builds the transaction: transfer to the recipient (90%) +\n transfer to the platform (10%), both carrying the uuid as a comment\n3. User signs it in their wallet (TON Connect)\n4. An indexer polls toncenter for incoming transfers to the platform\n wallet: comment == uuid AND amount == exact expected fee → confirmed\n5. Only after \"confirmed\" can the recipient open the text\n</code></pre>\n\n</div>\n\n\n\n<p>A few invariants I will die defending in code review:</p>\n\n<ul>\n<li>\n<strong>the server computes the price</strong>; the client's numbers are never trusted. The pricing formula lives in one shared module imported by both client (to display) and server (to verify) — they physically cannot diverge;</li>\n<li>\n<strong>every check happens before money moves</strong>: recipient exists? wallet connected? sender not blocked? text passed moderation? All before the transaction is built — because we <em>can't</em> refund, we never have the money;</li>\n<li>\n<strong>matching is idempotent</strong>: comment + exact amount; reprocessing the same transfer breaks nothing.</li>\n</ul>\n\n<p>The stack is boring, and that's a compliment: Bun + Fastify + bun:sqlite on the server, React + Vite + TON Connect UI on the client, TypeScript everywhere with shared types, deployed on Railway. No Postgres, no WebSocket, no Redis: SQLite on a volume and 7-second polling cover an MVP with room to spare. Every technology you don't adopt is a whole category of bugs you don't have.</p>\n\n<h2>\n\n\nRakes, paid for with real money\n</h2>\n\n<p>Debugging a payment pipeline has a particular property: unit tests are green, and the truth only comes out on real transfers.</p>\n\n<p><strong>Rake 1: toncenter returns comment = null.</strong> My uuid traveled in the jetton transfer's forward_payload, unit tests passed — and the production indexer just couldn't see the comment: the v3 API returned <code>comment: null</code>. Turns out you decode it yourself from the base64 BOC: open the cell, check op-code 0x0, read the string tail. Thirty minutes of \"the money left and the message didn't confirm\" panic — and a 15-line decoder.</p>\n\n<p><strong>Rake 2: raw vs friendly addresses.</strong> toncenter returns raw form (<code>0:hex</code>), wallets in TON Connect accept only friendly (<code>UQ…</code>). \"Wrong address format in message at index 0\" — on a live transfer, naturally. One Address.parse().toString() later it works, but you only learn this when a user (me) can't pay.</p>\n\n<p><strong>Rake 3: the webview caches your bundle forever.</strong> Telegram's webview loves serving a stale build after a deploy. A user reports a bug you fixed three versions ago. Fix: a version.json next to the bundle, client-side polling, and an \"update available\" banner.</p>\n\n<p><strong>Rake 4, my favorite: \"insufficient funds, 0.11 TON required\".</strong> Initially everything ran on USDT (a stablecoin — users think in dollars, seemed logical). Newcomers got 0.07 TON as a gas gift from a hot wallet so they could reply. Then a friend receives a message, money, the gift — taps \"reply\" — and the wallet demands 0.11 TON.</p>\n\n<p>The anatomy: a jetton transfer isn't \"send tokens\", it's a chain of contract calls, and the wallet requires the gas budget upfront, per transfer. A direct message = two transfers (recipient + fee) = 2 × 0.05 TON of prepaid budget. Only ~0.01 actually burns, the rest returns — but <strong>all of it must sit on the balance at send time</strong>. The 0.07 gift mathematically couldn't cover a reply.</p>\n\n<p>I bumped the gift, trimmed the budget… then realized I was treating symptoms. The disease was <strong>two-currency onboarding</strong>: users needed USDT for messages AND TON for gas. Normal for a crypto product. A funeral for a chat.</p>\n\n<h2>\n\n\nThe migration: throwing out the stablecoin\n</h2>\n\n<p>The fix looked like heresy: drop the stablecoin, price everything in native TON. The price is a constant 0.04 TON per character; dollars in the UI are just a reference line at the current rate.</p>\n\n<p>What happened after the migration:</p>\n\n<ul>\n<li>a TON transfer is a plain transaction with a value and a comment. No contract chains, no prepaid budgets: ~0.005 TON network fee, done;</li>\n<li>the exchange rate left the critical path: if the rate API dies, the dollar hints disappear — the money keeps working;</li>\n<li>amounts became round, matching became trivial, the indexer and client shed a third of their code;</li>\n<li>leaderboards became literally verifiable: the number in the top is the number on-chain, one to one.</li>\n</ul>\n\n<p>Volatility? Accepted deliberately: tickets are small, and if the rate moves several-fold, the per-character price is one constant to change. Fun fact: the original plan used the native coin; then I \"improved\" it to USDT; then reality reverted my improvement.</p>\n\n<h2>\n\n\nThe attention economy, in detail\n</h2>\n\n<p><strong>Threads in the public chat.</strong> I wanted replies — but how should they count? If replies boost the parent's rating ×2, nobody writes originals anymore, everyone parasites on hot roots. If they count for nothing, why reply? The answer: <strong>a reply is a payment to the author</strong>. 90% of a reply's price goes to the person you're replying to. Write something so good people <em>pay</em> to respond — your message literally earns. Ratings stay untouched: everyone pays for themselves, the crown belongs to a single message.</p>\n\n<p><strong>The newcomer gift.</strong> The first incoming direct message brings the recipient a bit of TON from the platform's hot wallet (limits: once per user, daily cap, total cap — so the wallet can't be drained). It's our money, not users' — non-custodial stays non-custodial. Why: an invited person can reply immediately without figuring out top-ups.</p>\n\n<p><strong>The profile as a billboard.</strong> If attention costs money, it can be resold: a user's profile has a \"card\" — a title, a line of text, a link. Write an expensive message → people open your profile → they see your ad. Links are https/t.me only, pre-moderated, reportable, and open behind a confirmation — so phishing can't ride on the bot's reputation.</p>\n\n<h2>\n\n\nSmall things that weren't small\n</h2>\n\n<ul>\n<li>\n<strong>Telegram webview on iOS</strong>: inputs under 16px auto-zoom; fullscreen collides with system buttons (fixed with Mini Apps 2.0 safe-area vars); the keyboard covers the composer (fixed by listening to viewportChanged + visualViewport and shrinking the layout);</li>\n<li>\n<strong>TON Connect sessions are per-device</strong>: connect on your phone, open the app on desktop — \"wallet not connected\". Catch it and quietly reopen the connect modal;</li>\n<li>\n<strong>Railway emails \"deployment crashed\" on every deploy</strong> if your process ignores SIGTERM: Bun exits with code 143, and a non-zero exit is a \"crash\" to Railway. Three lines of graceful shutdown — clean inbox;</li>\n<li>\n<strong>\\b in JS regex doesn't know Cyrillic</strong> — moderation with Russian word lists needs hand-built word boundaries;</li>\n<li>\n<strong>emoji must cost more</strong>: \"hello\" is 5 characters, but \"👨‍👩‍👧\" is 5 code points and one grapheme. Count graphemes via Intl.Segmenter, price any emoji as 4 characters — or you incentivize emoji-speak.</li>\n</ul>\n\n<h2>\n\n\nThe result\n</h2>\n\n<p>Three days from idea to an MVP with live transactions, another week of iterations with early users. 63 unit tests on everything that touches money. Zero custody of user funds, zero legal entities, ~$10/month of infrastructure.</p>\n\n<p>The first message in the product's history sold for $3: <em>\"Talk is cheap. This wasn't.\"</em></p>\n\n<p>If you want to poke it: <a href=\"https://t.me/centence_bot?start=src-devto\" rel=\"noopener noreferrer\">@centence_bot</a> — early users get a TON welcome gift, and I personally send paid messages to interesting newcomers because I'm testing push notifications.</p>\n\n<p>And a question for the room — an argument I keep having with myself: the indexer polls toncenter v3 every 7 seconds and at my volume it's flawless. Is there any point in webhooks/event streaming at small scale, or will polling forgive me for a long time yet? How do you confirm incoming TON payments in production?</p>","content":"Hi! I'm a mobile developer (Kotlin by day), and a couple of weeks ago I started a side project with a simple, slightly audacious idea: a chat where sending a message costs money — and the recipient gets paid. You write \"hi\", you pay. Someone writes to you, you earn.\nThis is the story of how \"what if attention literally had a price\" became a working Telegram Mini App with real on-chain transactions — and the rakes I stepped on along the way, from app store rules to the anatomy of TON gas. An LLM wrote the code with me, we argued about architecture together, but the debugging bills were mine — literally, since debugging happened on real money.\nThe idea: a market of expensive attention\nFree messages are worthless — in both senses. Spam, \"hey what's up\" from strangers, group chats with 10,000 unread messages. What if we flip it: every character costs money?\nThe economics get interesting fast:\n- you can message any stranger — but you'll pay them. A message with money inside always gets opened;\n- the public chat becomes an auction of wit: the priciest message of the week wears the crown 👑 and stays pinned;\n- spam dies on its own: spamming at $0.05 per character is a fast way to go broke.\nI called it Centence. Between the idea and the product stood two walls.\nWall one: the app stores\nA mobile developer's first reflex is a native app. It's also the last one: on iOS and Android any payment for \"digital content\" must go through in-app purchases — 30% commission and rules that treat P2P money transfers between users as a minefield, and crypto as a minefield in fog.\nOption two: a Telegram Mini App. But there's a rule there too — digital goods and services inside mini apps must be sold for Telegram Stars. Paying for \"a message\" with Stars, with the platform taking a cut, killed the whole point: money must go to the recipient, directly.\nThen came the pleasant realization: the Stars rule covers buying digital goods from the app. If the app sells nothing and never touches money — if it's just an interface through which one person transfers money from their wallet to another person's wallet — that's not a purchase. That's a transfer.\nWall two: the legal entity\nThe moment a service accepts users' money into its own account (even \"for a second\", even \"just to forward it\"), it becomes a money operator. That means licenses, KYC/AML, a legal entity — everything that kills a side project on day one.\nHence the architectural principle that everything else obeys:\nThe platform never touches user funds. No app balance, no deposits, no withdrawals, no escrow, no refunds. The server holds no private keys and is physically incapable of moving anyone's money.\nA payment goes from the sender's wallet straight to the recipient's wallet via TON Connect. My fee (10% on direct messages) is a separate output of the same transaction, to my own wallet. If the server gets hacked — there's no money in it. None.\nLegal bonus: since we're not a money operator, no licenses and no legal entity needed. Taxes on the fee are ordinary personal income.\nThe architecture: matching a transaction to a message\nThe central technical problem of a non-custodial design: the server must figure out that a specific on-chain transaction pays for a specific message — while the transaction is signed by the user in their own wallet, and until the network confirms it, no \"payment\" exists.\nThe scheme:\n1. Client sends the text to the server → server computes the price,\nstores {uuid, text, amounts, status: pending}\n2. Client builds the transaction: transfer to the recipient (90%) +\ntransfer to the platform (10%), both carrying the uuid as a comment\n3. User signs it in their wallet (TON Connect)\n4. An indexer polls toncenter for incoming transfers to the platform\nwallet: comment == uuid AND amount == exact expected fee → confirmed\n5. Only after \"confirmed\" can the recipient open the text\nA few invariants I will die defending in code review:\n- the server computes the price; the client's numbers are never trusted. The pricing formula lives in one shared module imported by both client (to display) and server (to verify) — they physically cannot diverge;\n- every check happens before money moves: recipient exists? wallet connected? sender not blocked? text passed moderation? All before the transaction is built — because we can't refund, we never have the money;\n- matching is idempotent: comment + exact amount; reprocessing the same transfer breaks nothing.\nThe stack is boring, and that's a compliment: Bun + Fastify + bun:sqlite on the server, React + Vite + TON Connect UI on the client, TypeScript everywhere with shared types, deployed on Railway. No Postgres, no WebSocket, no Redis: SQLite on a volume and 7-second polling cover an MVP with room to spare. Every technology you don't adopt is a whole category of bugs you don't have.\nRakes, paid for with real money\nDebugging a payment pipeline has a particular property: unit tests are green, and the truth only comes out on real transfers.\nRake 1: toncenter returns comment = null. My uuid traveled in the jetton transfer's forward_payload, unit tests passed — and the production indexer just couldn't see the comment: the v3 API returned comment: null\n. Turns out you decode it yourself from the base64 BOC: open the cell, check op-code 0x0, read the string tail. Thirty minutes of \"the money left and the message didn't confirm\" panic — and a 15-line decoder.\nRake 2: raw vs friendly addresses. toncenter returns raw form (0:hex\n), wallets in TON Connect accept only friendly (UQ…\n). \"Wrong address format in message at index 0\" — on a live transfer, naturally. One Address.parse().toString() later it works, but you only learn this when a user (me) can't pay.\nRake 3: the webview caches your bundle forever. Telegram's webview loves serving a stale build after a deploy. A user reports a bug you fixed three versions ago. Fix: a version.json next to the bundle, client-side polling, and an \"update available\" banner.\nRake 4, my favorite: \"insufficient funds, 0.11 TON required\". Initially everything ran on USDT (a stablecoin — users think in dollars, seemed logical). Newcomers got 0.07 TON as a gas gift from a hot wallet so they could reply. Then a friend receives a message, money, the gift — taps \"reply\" — and the wallet demands 0.11 TON.\nThe anatomy: a jetton transfer isn't \"send tokens\", it's a chain of contract calls, and the wallet requires the gas budget upfront, per transfer. A direct message = two transfers (recipient + fee) = 2 × 0.05 TON of prepaid budget. Only ~0.01 actually burns, the rest returns — but all of it must sit on the balance at send time. The 0.07 gift mathematically couldn't cover a reply.\nI bumped the gift, trimmed the budget… then realized I was treating symptoms. The disease was two-currency onboarding: users needed USDT for messages AND TON for gas. Normal for a crypto product. A funeral for a chat.\nThe migration: throwing out the stablecoin\nThe fix looked like heresy: drop the stablecoin, price everything in native TON. The price is a constant 0.04 TON per character; dollars in the UI are just a reference line at the current rate.\nWhat happened after the migration:\n- a TON transfer is a plain transaction with a value and a comment. No contract chains, no prepaid budgets: ~0.005 TON network fee, done;\n- the exchange rate left the critical path: if the rate API dies, the dollar hints disappear — the money keeps working;\n- amounts became round, matching became trivial, the indexer and client shed a third of their code;\n- leaderboards became literally verifiable: the number in the top is the number on-chain, one to one.\nVolatility? Accepted deliberately: tickets are small, and if the rate moves several-fold, the per-character price is one constant to change. Fun fact: the original plan used the native coin; then I \"improved\" it to USDT; then reality reverted my improvement.\nThe attention economy, in detail\nThreads in the public chat. I wanted replies — but how should they count? If replies boost the parent's rating ×2, nobody writes originals anymore, everyone parasites on hot roots. If they count for nothing, why reply? The answer: a reply is a payment to the author. 90% of a reply's price goes to the person you're replying to. Write something so good people pay to respond — your message literally earns. Ratings stay untouched: everyone pays for themselves, the crown belongs to a single message.\nThe newcomer gift. The first incoming direct message brings the recipient a bit of TON from the platform's hot wallet (limits: once per user, daily cap, total cap — so the wallet can't be drained). It's our money, not users' — non-custodial stays non-custodial. Why: an invited person can reply immediately without figuring out top-ups.\nThe profile as a billboard. If attention costs money, it can be resold: a user's profile has a \"card\" — a title, a line of text, a link. Write an expensive message → people open your profile → they see your ad. Links are https/t.me only, pre-moderated, reportable, and open behind a confirmation — so phishing can't ride on the bot's reputation.\nSmall things that weren't small\n- Telegram webview on iOS: inputs under 16px auto-zoom; fullscreen collides with system buttons (fixed with Mini Apps 2.0 safe-area vars); the keyboard covers the composer (fixed by listening to viewportChanged + visualViewport and shrinking the layout);\n- TON Connect sessions are per-device: connect on your phone, open the app on desktop — \"wallet not connected\". Catch it and quietly reopen the connect modal;\n- Railway emails \"deployment crashed\" on every deploy if your process ignores SIGTERM: Bun exits with code 143, and a non-zero exit is a \"crash\" to Railway. Three lines of graceful shutdown — clean inbox;\n- \\b in JS regex doesn't know Cyrillic — moderation with Russian word lists needs hand-built word boundaries;\n- emoji must cost more: \"hello\" is 5 characters, but \"👨👩👧\" is 5 code points and one grapheme. Count graphemes via Intl.Segmenter, price any emoji as 4 characters — or you incentivize emoji-speak.\nThe result\nThree days from idea to an MVP with live transactions, another week of iterations with early users. 63 unit tests on everything that touches money. Zero custody of user funds, zero legal entities, ~$10/month of infrastructure.\nThe first message in the product's history sold for $3: \"Talk is cheap. This wasn't.\"\nIf you want to poke it: @centence_bot — early users get a TON welcome gift, and I personally send paid messages to interesting newcomers because I'm testing push notifications.\nAnd a question for the room — an argument I keep having with myself: the indexer polls toncenter v3 every 7 seconds and at my volume it's flawless. Is there any point in webhooks/event streaming at small scale, or will polling forgive me for a long time yet? How do you confirm incoming TON payments in production?\nTop comments (0)","published":"Sat, 12 Sep 2026 23:07:55 +0000","author":"Антон Петровский","guid":"https://dev.to/__e3096294/i-built-a-chat-where-every-word-costs-money-on-ton-solo-no-legal-entity-1f3f","created_at":"2026-09-13T01:36:17.352356","last_synchronized":"2026-09-13T01:36:17.352356","sentiment":{"sentiment":"Positive","score":0.9964,"details":{"neg":0.053,"neu":0.866,"pos":0.081,"compound":0.9964}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"‘I Just Feel Ripped Off’: A Week of Users Asking What They Pay For","link":"https://dev.to/theaidownside/i-just-feel-ripped-off-a-week-of-users-asking-what-they-pay-for-5em6","description":"<p>There is a particular kind of complaint that gets louder as a market matures. Not “this is broken” — that is the noise of a new product — but “wait, what am I actually paying for?” That is the sound of customers who have done the sum. This week, across the AI tools people pay for daily, that is the sound that carried.</p>\n\n<p>None of it is a single scandal. It is a mood: ads turning up inside plans that used to be clean, “unlimited” offers with a quiet expiry date, top-tier subscriptions that somehow feel slower than the cheap ones, and the steady suspicion that the thing you rented last month has been swapped for something worse. Individually, each is a shrug. Together, they read as a slow renegotiation of the deal — in the house’s favour, one default at a time.</p>\n\n<p><strong>Quotes sourced from: Reddit and Hacker News.</strong> As ever, these are verified against the live posts and linked in full at the foot of the piece; we quote real users to show the texture of the frustration, not to pretend a forum is a poll.</p>\n\n<h2>\n\n\nThe ads arrive inside the thing you paid for\n</h2>\n\n<p>The sharpest version came from a paying ChatGPT Go subscriber, writing on r/OpenAI, who had been broadly content until advertising started appearing in the chat itself. “When I subscribed to Go, there were no ads in my conversations,” wrote u/Loganh1976 on 26 August. “I was never expecting to suddenly find advertising inserted into the conversations months later.” The kicker was strategic, not just grumpy: the ads, they said, had done what nothing else had, and pushed them to start trying Claude, Gemini and Grok for the first time.</p>\n\n<blockquote>\n<p>Moan of the day: “I genuinely think putting intrusive advertising into a paid plan is one of the worst strategic decisions OpenAI could have made.” — u/Loganh1976, r/OpenAI</p>\n</blockquote>\n\n<p>The mechanism here is the oldest one in subscriptions: a plan sold as the clean, paid alternative gets monetised a second time once you are inside it. We have watched the pricing on these products drift from a simple subscription toward <a href=\"https://theaidownside.com/posts/your-flat-ai-subscription-is-becoming-a-meter.html\" rel=\"noopener noreferrer\">something closer to a meter</a>, and the arrival of ads in a paid tier — the same shift Europe saw when <a href=\"https://theaidownside.com/posts/chatgpt-ads-come-to-europe.html\" rel=\"noopener noreferrer\">ChatGPT ads reached the free plan</a> — is the same instinct climbing up the price ladder.</p>\n\n<h2>\n\n\n“Unlimited”, until it isn’t\n</h2>\n\n<p>Over on r/cursor, the coding-tool crowd were doing the arithmetic on their own plans. One long-time subscriber captured the churn precisely: “I have spent most of this year exclusively with Claude,” wrote u/-AMARYANA- on 26 August. “I just feel ripped off at this point. I see the value is dropping each month as other models close the gap.” Another, u/General-History-5917, was staring at the practical version of the problem — the annual Cursor Pro plan with “unlimited auto mode” ending, forcing a choice about which more-metered tier to move to next.</p>\n\n<p>“Unlimited” that expires is not a bug; it is a launch tactic reaching its scheduled end. But it lands as a downgrade, because it is one. The pattern — introduce a generous flat allowance to win the habit, then convert it to usage-based pricing once the habit is formed — is exactly the shift that leaves heavy users feeling the ground move under a plan they thought they understood.</p>\n\n<h2>\n\n\nPaying the most, waiting the longest\n</h2>\n\n<p>The oddest complaints came from the people spending the most. On r/OpenAI, a subscriber to the top “20x” plan wrote that the premium tier was the frustrating one: “I am seeing a lot of hallucination and scope drift, and extremely slow execution,” said u/SweatyActuator2119 on 26 August, describing new models that take forever and “constantly over-engineer to the point where scope is an afterthought.” The instinct that expensive should mean better runs straight into a product where the priciest, most “thoughtful” modes are also the slowest.</p>\n\n<p>Claude users had their own version of paying-for-friction. One, posting on r/ClaudeAI, warned that even sticking with a trusted older model doesn’t save you: pick “the good old Opus 4.6” for careful work, wrote u/SemiMagnum, and the main model may still hand subtasks to “verbose and token-consuming Opus 5 subagents ruining your work and burning the limits.” You choose the model you trust; the system spends your allowance on the one you didn’t.</p>\n\n<h2>\n\n\nThe model you rented got quietly worse\n</h2>\n\n<p>A recurring suspicion this week was that older models are being throttled to nudge people onto newer ones. A Gemini developer laid it out with numbers, on r/GeminiAI: short requests that “used to take between 4 and 15 seconds” were now, “over the past week,” taking “50s or more.” Their read, offered as a guess rather than a claim: “Google reduces capacity for older models to push people to move to the latest.”</p>\n\n<p>We can’t confirm the motive, and neither can they — that is the whole problem with a service you can’t inspect. But the lived effect is real and measurable from the user’s chair, and it is the same complaint we keep hearing: <a href=\"https://theaidownside.com/posts/claudes-rate-limits-are-still-confusing.html\" rel=\"noopener noreferrer\">the terms of what you’re getting keep shifting</a> without a memo, and always in the direction of the upgrade.</p>\n\n<h2>\n\n\nWhat did I actually get billed for?\n</h2>\n\n<p>Some of the week’s unease was simpler still: people who genuinely could not tell what they had paid for. An “Ask HN” thread was titled, plaintively, “What Did Anthropic Bill Me For?” — its author, u/OhMeadhbh, reduced to guessing whether a charge had become credits that the interface simply wasn’t showing. When the bill is a mystery and the fix is to ask the chatbot how to reach a human, the transparency problem has stopped being a detail.</p>\n\n<p>And the upsell, when it is visible, grates. A Perplexity subscriber on r/perplexity_ai was blunt: “Perplexity is desperately using underhanded techniques to upsell, plus has no customer service. Not a fan anymore.” Whether or not you share the verdict, the underlying complaint — that the product is working harder to sell you the next tier than to answer the question — is the through-line of the whole week. Even the once-cheap options drew the same sigh: a poster on r/DeepSeek signed off hunting for a workaround “for someone out there yearning for Cheapseek prices” — a small joke that only lands because the famously cheap option no longer feels it.</p>\n\n<h2>\n\n\nPaying for a tool that won’t follow the rules\n</h2>\n\n<p>The most basic value complaint is the one where the tool won’t do as it’s told. On Hacker News, in a thread on what large language models are still bad at, one developer described the daily reality of paying for a top model and repeating yourself anyway: “I’ve got a modest sized CLAUDE.md containing some simple rules to follow. Things to always do, things to never do,” wrote u/jgb1984. “Not a day goes by where Claude Opus violates one or several of the instructions.” It is a small thing and a fundamental one at once: the pitch of a premium assistant is that it saves you effort, and re-issuing the same rules every session is effort the product promised to remove.</p>\n\n<h2>\n\n\nThe moves, in one place\n</h2>\n\n<p>Pull the week’s gripes apart and the same handful of manoeuvres keep showing up. None is illegal; none is even unusual. It is the accumulation, and the quietness, that wears people down:</p>\n\n<ul>\n<li>\n<strong>Add a revenue stream to a plan you already sold</strong> — ads inside a paid tier that used to have none.</li>\n<li>\n<strong>Launch generous, then meter</strong> — an “unlimited” allowance that wins the habit, then expires into usage-based pricing.</li>\n<li>\n<strong>Charge most for the slowest</strong> — premium tiers whose “thinking” modes deliver latency and over-engineering rather than a clearer win.</li>\n<li>\n<strong>Let the system overrule your choice</strong> — pick a trusted model and watch subtasks get handed to a costlier one that burns your limits.</li>\n<li>\n<strong>Quietly slow the old thing</strong> — degrade older models so the upgrade looks compelling, without saying you’ve done it.</li>\n<li>\n<strong>Make the bill unreadable</strong> — credits, tiers and charges opaque enough that you can’t easily tell what you paid for.</li>\n</ul>\n\n<h2>\n\n\nThe switch that used to feel impossible\n</h2>\n\n<p>Here is the thread that ties the week together, and it should trouble the incumbents more than any single gripe: the people complaining are not, mostly, threatening to quit in a rage. They are calmly shopping. The ChatGPT Go subscriber annoyed by ads said the ads had done what nothing else managed — made them start trying Claude, Gemini and Grok. The Cursor loyalist who felt “ripped off” was weighing where next month’s twenty dollars should go. The frustrated top-tier subscriber was openly asking whether to switch to Claude’s equivalent plan or drop back to a cheaper open-weight setup and actually get some work done.</p>\n\n<p>For years the moat around these products was inertia. The models were different enough, and the effort of moving accounts, prompts and habits high enough, that grumbling rarely hardened into leaving. That moat is draining. As the tools converge on quality and nearly everyone ships a harness that will happily run someone else’s model, the cost of trying the competitor drops to an idle afternoon. When switching is that easy, “what am I paying for?” stops being a rhetorical sigh at the end of a bad week and becomes a live question with a cheaper answer sitting one browser tab away. The companies still have the better demos. What they are visibly losing, this week, is the benefit of the doubt.</p>\n\n<h2>\n\n\nThe fair version, and what to do about it\n</h2>\n\n<p>To be fair, because it matters: serving these tools is genuinely expensive, heavy users really do cost more than they pay, and some of what reads as a slowdown is ordinary variance in models and infrastructure that change constantly. A company adjusting an unsustainable “unlimited” offer is not a conspiracy, and a forum full of the annoyed is not a representative sample. Concede all of it.</p>\n\n<p>What the concession doesn’t buy is the quietness. The recurring injury this week wasn’t that prices exist; it was that the deal keeps changing under people who are still paying the same amount — ads added, limits tightened, models swapped, bills obscured — without anyone being told. So do the unglamorous things. Check what your plan guarantees against what it merely implies. Note your renewal date, and treat any “unlimited” as a countdown. Keep your own tally of the slow days and the vanished features, because a flat fee is counting on you not to. And remember the one lever you always hold: several of this week’s posters, for the first time, were not cancelling in a huff — they were calmly opening a competitor to compare. That is the sound a market makes when it stops taking the deal on trust — and it is a far more dangerous sound, for the companies, than any amount of shouting.</p>\n\n\n\n\n<p><em>Originally published at <a href=\"https://theaidownside.com/posts/voices-what-am-i-paying-for.html\" rel=\"noopener noreferrer\">theaidownside.com</a> — evidence-first reporting on the costs and trade-offs behind AI products.</em></p>","content":"There is a particular kind of complaint that gets louder as a market matures. Not “this is broken” — that is the noise of a new product — but “wait, what am I actually paying for?” That is the sound of customers who have done the sum. This week, across the AI tools people pay for daily, that is the sound that carried.\nNone of it is a single scandal. It is a mood: ads turning up inside plans that used to be clean, “unlimited” offers with a quiet expiry date, top-tier subscriptions that somehow feel slower than the cheap ones, and the steady suspicion that the thing you rented last month has been swapped for something worse. Individually, each is a shrug. Together, they read as a slow renegotiation of the deal — in the house’s favour, one default at a time.\nQuotes sourced from: Reddit and Hacker News. As ever, these are verified against the live posts and linked in full at the foot of the piece; we quote real users to show the texture of the frustration, not to pretend a forum is a poll.\nThe ads arrive inside the thing you paid for\nThe sharpest version came from a paying ChatGPT Go subscriber, writing on r/OpenAI, who had been broadly content until advertising started appearing in the chat itself. “When I subscribed to Go, there were no ads in my conversations,” wrote u/Loganh1976 on 26 August. “I was never expecting to suddenly find advertising inserted into the conversations months later.” The kicker was strategic, not just grumpy: the ads, they said, had done what nothing else had, and pushed them to start trying Claude, Gemini and Grok for the first time.\nMoan of the day: “I genuinely think putting intrusive advertising into a paid plan is one of the worst strategic decisions OpenAI could have made.” — u/Loganh1976, r/OpenAI\nThe mechanism here is the oldest one in subscriptions: a plan sold as the clean, paid alternative gets monetised a second time once you are inside it. We have watched the pricing on these products drift from a simple subscription toward something closer to a meter, and the arrival of ads in a paid tier — the same shift Europe saw when ChatGPT ads reached the free plan — is the same instinct climbing up the price ladder.\n“Unlimited”, until it isn’t\nOver on r/cursor, the coding-tool crowd were doing the arithmetic on their own plans. One long-time subscriber captured the churn precisely: “I have spent most of this year exclusively with Claude,” wrote u/-AMARYANA- on 26 August. “I just feel ripped off at this point. I see the value is dropping each month as other models close the gap.” Another, u/General-History-5917, was staring at the practical version of the problem — the annual Cursor Pro plan with “unlimited auto mode” ending, forcing a choice about which more-metered tier to move to next.\n“Unlimited” that expires is not a bug; it is a launch tactic reaching its scheduled end. But it lands as a downgrade, because it is one. The pattern — introduce a generous flat allowance to win the habit, then convert it to usage-based pricing once the habit is formed — is exactly the shift that leaves heavy users feeling the ground move under a plan they thought they understood.\nPaying the most, waiting the longest\nThe oddest complaints came from the people spending the most. On r/OpenAI, a subscriber to the top “20x” plan wrote that the premium tier was the frustrating one: “I am seeing a lot of hallucination and scope drift, and extremely slow execution,” said u/SweatyActuator2119 on 26 August, describing new models that take forever and “constantly over-engineer to the point where scope is an afterthought.” The instinct that expensive should mean better runs straight into a product where the priciest, most “thoughtful” modes are also the slowest.\nClaude users had their own version of paying-for-friction. One, posting on r/ClaudeAI, warned that even sticking with a trusted older model doesn’t save you: pick “the good old Opus 4.6” for careful work, wrote u/SemiMagnum, and the main model may still hand subtasks to “verbose and token-consuming Opus 5 subagents ruining your work and burning the limits.” You choose the model you trust; the system spends your allowance on the one you didn’t.\nThe model you rented got quietly worse\nA recurring suspicion this week was that older models are being throttled to nudge people onto newer ones. A Gemini developer laid it out with numbers, on r/GeminiAI: short requests that “used to take between 4 and 15 seconds” were now, “over the past week,” taking “50s or more.” Their read, offered as a guess rather than a claim: “Google reduces capacity for older models to push people to move to the latest.”\nWe can’t confirm the motive, and neither can they — that is the whole problem with a service you can’t inspect. But the lived effect is real and measurable from the user’s chair, and it is the same complaint we keep hearing: the terms of what you’re getting keep shifting without a memo, and always in the direction of the upgrade.\nWhat did I actually get billed for?\nSome of the week’s unease was simpler still: people who genuinely could not tell what they had paid for. An “Ask HN” thread was titled, plaintively, “What Did Anthropic Bill Me For?” — its author, u/OhMeadhbh, reduced to guessing whether a charge had become credits that the interface simply wasn’t showing. When the bill is a mystery and the fix is to ask the chatbot how to reach a human, the transparency problem has stopped being a detail.\nAnd the upsell, when it is visible, grates. A Perplexity subscriber on r/perplexity_ai was blunt: “Perplexity is desperately using underhanded techniques to upsell, plus has no customer service. Not a fan anymore.” Whether or not you share the verdict, the underlying complaint — that the product is working harder to sell you the next tier than to answer the question — is the through-line of the whole week. Even the once-cheap options drew the same sigh: a poster on r/DeepSeek signed off hunting for a workaround “for someone out there yearning for Cheapseek prices” — a small joke that only lands because the famously cheap option no longer feels it.\nPaying for a tool that won’t follow the rules\nThe most basic value complaint is the one where the tool won’t do as it’s told. On Hacker News, in a thread on what large language models are still bad at, one developer described the daily reality of paying for a top model and repeating yourself anyway: “I’ve got a modest sized CLAUDE.md containing some simple rules to follow. Things to always do, things to never do,” wrote u/jgb1984. “Not a day goes by where Claude Opus violates one or several of the instructions.” It is a small thing and a fundamental one at once: the pitch of a premium assistant is that it saves you effort, and re-issuing the same rules every session is effort the product promised to remove.\nThe moves, in one place\nPull the week’s gripes apart and the same handful of manoeuvres keep showing up. None is illegal; none is even unusual. It is the accumulation, and the quietness, that wears people down:\n- Add a revenue stream to a plan you already sold — ads inside a paid tier that used to have none.\n- Launch generous, then meter — an “unlimited” allowance that wins the habit, then expires into usage-based pricing.\n- Charge most for the slowest — premium tiers whose “thinking” modes deliver latency and over-engineering rather than a clearer win.\n- Let the system overrule your choice — pick a trusted model and watch subtasks get handed to a costlier one that burns your limits.\n- Quietly slow the old thing — degrade older models so the upgrade looks compelling, without saying you’ve done it.\n- Make the bill unreadable — credits, tiers and charges opaque enough that you can’t easily tell what you paid for.\nThe switch that used to feel impossible\nHere is the thread that ties the week together, and it should trouble the incumbents more than any single gripe: the people complaining are not, mostly, threatening to quit in a rage. They are calmly shopping. The ChatGPT Go subscriber annoyed by ads said the ads had done what nothing else managed — made them start trying Claude, Gemini and Grok. The Cursor loyalist who felt “ripped off” was weighing where next month’s twenty dollars should go. The frustrated top-tier subscriber was openly asking whether to switch to Claude’s equivalent plan or drop back to a cheaper open-weight setup and actually get some work done.\nFor years the moat around these products was inertia. The models were different enough, and the effort of moving accounts, prompts and habits high enough, that grumbling rarely hardened into leaving. That moat is draining. As the tools converge on quality and nearly everyone ships a harness that will happily run someone else’s model, the cost of trying the competitor drops to an idle afternoon. When switching is that easy, “what am I paying for?” stops being a rhetorical sigh at the end of a bad week and becomes a live question with a cheaper answer sitting one browser tab away. The companies still have the better demos. What they are visibly losing, this week, is the benefit of the doubt.\nThe fair version, and what to do about it\nTo be fair, because it matters: serving these tools is genuinely expensive, heavy users really do cost more than they pay, and some of what reads as a slowdown is ordinary variance in models and infrastructure that change constantly. A company adjusting an unsustainable “unlimited” offer is not a conspiracy, and a forum full of the annoyed is not a representative sample. Concede all of it.\nWhat the concession doesn’t buy is the quietness. The recurring injury this week wasn’t that prices exist; it was that the deal keeps changing under people who are still paying the same amount — ads added, limits tightened, models swapped, bills obscured — without anyone being told. So do the unglamorous things. Check what your plan guarantees against what it merely implies. Note your renewal date, and treat any “unlimited” as a countdown. Keep your own tally of the slow days and the vanished features, because a flat fee is counting on you not to. And remember the one lever you always hold: several of this week’s posters, for the first time, were not cancelling in a huff — they were calmly opening a competitor to compare. That is the sound a market makes when it stops taking the deal on trust — and it is a far more dangerous sound, for the companies, than any amount of shouting.\nOriginally published at theaidownside.com — evidence-first reporting on the costs and trade-offs behind AI products.\nTop comments (0)","published":"Sat, 12 Sep 2026 23:05:14 +0000","author":"The AI Downside","guid":"https://dev.to/theaidownside/i-just-feel-ripped-off-a-week-of-users-asking-what-they-pay-for-5em6","created_at":"2026-09-13T01:36:17.352356","last_synchronized":"2026-09-13T01:36:17.352356","sentiment":{"sentiment":"Positive","score":0.9277,"details":{"neg":0.072,"neu":0.849,"pos":0.08,"compound":0.9277}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"I built an hourly newspaper for e-ink (and turned the pipeline into an MCP server)","link":"https://dev.to/jshelley/i-built-an-hourly-newspaper-for-e-ink-and-turned-the-pipeline-into-an-mcp-server-1ebg","description":"<p>Every hour, a Lambda pulls about 120 items from ~100 RSS feeds, asks a small model to act as an editor, and typesets the result as a 480x800 one-bit page for the Xteink X4 e-ink reader on my desk. The same pipeline now produces a multi-page \"Morning Paper\" PDF for reMarkable and Kindle, a TRMNL plugin, and an MCP server that agents can call. This is what it does, what broke, and what I learned about \"news for agents\". Everything below is live at <a href=\"https://briefing-service.wholemind.workers.dev\" rel=\"noopener noreferrer\">https://briefing-service.wholemind.workers.dev</a>.</p>\n\n<h2>\n\n\nThe pipeline\n</h2>\n\n<ol>\n<li>\n<strong>Collect.</strong> feedparser over ~100 feeds per briefing (AI, world, US, finance, sports, soccer, frontier-lab blogs). Each item keeps id, source, title, link, canonical link (utm and friends stripped), published, a 320-char summary, and a per-feed weight.</li>\n<li>\n<strong>Dedupe.</strong> Exact duplicates collapse by canonical URL. My first near-duplicate rule (title+summary word-shingle Jaccard &gt;= 0.5) measured zero merges on 480 live candidates, and only five false pairs even at 0.2: outlets rewrite wire copy, and RSS summaries share boilerplate, not story text. What works is headline entity overlap: adjacent capitalised words form one entity (\"Wall Street\", \"Joao Pedro\"), punctuation and hyphens end a phrase, Title-Case outlets are filtered against their own summary, and two headlines from different outlets merge on three shared entities or two non-generic ones. On the same 429 news items that gives nine merges, all the same story on inspection. The survivor lists the other outlets in <code>also_in</code>; the collapsed records are published too, with <code>duplicate_of</code>, the rule and a score, so anyone doing provenance work can audit the merges.</li>\n<li>\n<strong>Rank.</strong> Claude Haiku 4.5 on Bedrock gets the candidate list (id, source, age, title, summary) and a persona, and must call a <code>publish_briefing</code> tool with a lead, N stories, a research section, one-sentence summaries, a three-to-five-sentence detail passage, a why-it-matters line and key points. Facts must come from candidate text; the tool schema is the guardrail. A non-LLM fallback ranker (weight, recency, two-per-outlet cap) runs if the model call fails, so the device never shows a blank page.</li>\n<li>\n<strong>Render.</strong> Pillow draws a masthead, lead, numbered stories and a research strip into 480x800, dithers to 1-bit, writes BMPs the reader's firmware can page through, plus a wide 800x480 variant for TRMNL-class panels and a PDF for e-readers.</li>\n<li>\n<strong>Publish.</strong> S3 + CloudFront, a tiny manifest with a stamp that changes only when the editor's picks change, so the device never re-downloads unchanged pages.</li>\n</ol>\n\n<h2>\n\n\nWhat surprised me\n</h2>\n\n<ul>\n<li>\n<strong>The editor is the product.</strong> The rendering is fun, but the thing people react to is \"lead + why it matters\" over a hundred sources with duplicates merged. That is why I exposed it as an API and an MCP server rather than keeping it a device toy.</li>\n<li>\n<strong>Agents are terrible customers so far.</strong> Listing the MCP server in the official registry, Glama, Smithery and a few awesome-lists produced steady traffic in a day: scanners, auditors, and health checks. Thirteen MCP calls, zero humans. If you are building \"for agents\", expect the first wave to be bots evaluating you.</li>\n<li>\n<strong>E-ink people want files, not APIs.</strong> reMarkable and Kindle owners asked for a PDF in their library each morning, so that exists (rmapi push, Send-to-Kindle email), free for the first ten.</li>\n</ul>\n\n<h2>\n\n\nTry it\n</h2>\n\n<ul>\n<li>Read: the front pages and JSON are free, 25 API/MCP calls per IP per day: <a href=\"https://briefing-service.wholemind.workers.dev\" rel=\"noopener noreferrer\">https://briefing-service.wholemind.workers.dev</a>\n</li>\n<li>Provenance work: <a href=\"https://briefing-service.wholemind.workers.dev/v1/briefings/ai/candidates\" rel=\"noopener noreferrer\">https://briefing-service.wholemind.workers.dev/v1/briefings/ai/candidates</a> is the full hourly candidate set, free, no key.</li>\n<li>Agents: MCP endpoint at <a href=\"https://briefing-service.wholemind.workers.dev/mcp\" rel=\"noopener noreferrer\">https://briefing-service.wholemind.workers.dev/mcp</a> (streamable HTTP), tools <code>list_briefings</code>, <code>get_briefing</code>, <code>get_page</code>, <code>render_briefing</code>. Installable: <code>npx -y github:jshelley/briefing-mcp</code>.</li>\n<li>Paid: $9/month Reader key, x402 per call (USDC on Base), $49/month Team Briefing over your own feeds.</li>\n</ul>\n\n<p>I would like to hear what a \"news\" tool should return to your agent: full JSON, a short digest, or the rendered page.</p>","content":"Every hour, a Lambda pulls about 120 items from ~100 RSS feeds, asks a small model to act as an editor, and typesets the result as a 480x800 one-bit page for the Xteink X4 e-ink reader on my desk. The same pipeline now produces a multi-page \"Morning Paper\" PDF for reMarkable and Kindle, a TRMNL plugin, and an MCP server that agents can call. This is what it does, what broke, and what I learned about \"news for agents\". Everything below is live at https://briefing-service.wholemind.workers.dev.\nThe pipeline\n- Collect. feedparser over ~100 feeds per briefing (AI, world, US, finance, sports, soccer, frontier-lab blogs). Each item keeps id, source, title, link, canonical link (utm and friends stripped), published, a 320-char summary, and a per-feed weight.\n-\nDedupe. Exact duplicates collapse by canonical URL. My first near-duplicate rule (title+summary word-shingle Jaccard >= 0.5) measured zero merges on 480 live candidates, and only five false pairs even at 0.2: outlets rewrite wire copy, and RSS summaries share boilerplate, not story text. What works is headline entity overlap: adjacent capitalised words form one entity (\"Wall Street\", \"Joao Pedro\"), punctuation and hyphens end a phrase, Title-Case outlets are filtered against their own summary, and two headlines from different outlets merge on three shared entities or two non-generic ones. On the same 429 news items that gives nine merges, all the same story on inspection. The survivor lists the other outlets in\nalso_in\n; the collapsed records are published too, withduplicate_of\n, the rule and a score, so anyone doing provenance work can audit the merges. -\nRank. Claude Haiku 4.5 on Bedrock gets the candidate list (id, source, age, title, summary) and a persona, and must call a\npublish_briefing\ntool with a lead, N stories, a research section, one-sentence summaries, a three-to-five-sentence detail passage, a why-it-matters line and key points. Facts must come from candidate text; the tool schema is the guardrail. A non-LLM fallback ranker (weight, recency, two-per-outlet cap) runs if the model call fails, so the device never shows a blank page. - Render. Pillow draws a masthead, lead, numbered stories and a research strip into 480x800, dithers to 1-bit, writes BMPs the reader's firmware can page through, plus a wide 800x480 variant for TRMNL-class panels and a PDF for e-readers.\n- Publish. S3 + CloudFront, a tiny manifest with a stamp that changes only when the editor's picks change, so the device never re-downloads unchanged pages.\nWhat surprised me\n- The editor is the product. The rendering is fun, but the thing people react to is \"lead + why it matters\" over a hundred sources with duplicates merged. That is why I exposed it as an API and an MCP server rather than keeping it a device toy.\n- Agents are terrible customers so far. Listing the MCP server in the official registry, Glama, Smithery and a few awesome-lists produced steady traffic in a day: scanners, auditors, and health checks. Thirteen MCP calls, zero humans. If you are building \"for agents\", expect the first wave to be bots evaluating you.\n- E-ink people want files, not APIs. reMarkable and Kindle owners asked for a PDF in their library each morning, so that exists (rmapi push, Send-to-Kindle email), free for the first ten.\nTry it\n- Read: the front pages and JSON are free, 25 API/MCP calls per IP per day: https://briefing-service.wholemind.workers.dev\n- Provenance work: https://briefing-service.wholemind.workers.dev/v1/briefings/ai/candidates is the full hourly candidate set, free, no key.\n- Agents: MCP endpoint at https://briefing-service.wholemind.workers.dev/mcp (streamable HTTP), tools\nlist_briefings\n,get_briefing\n,get_page\n,render_briefing\n. Installable:npx -y github:jshelley/briefing-mcp\n. - Paid: $9/month Reader key, x402 per call (USDC on Base), $49/month Team Briefing over your own feeds.\nI would like to hear what a \"news\" tool should return to your agent: full JSON, a short digest, or the rendered page.\nTop comments (0)","published":"Sat, 12 Sep 2026 22:54:03 +0000","author":"Jordan Shelley","guid":"https://dev.to/jshelley/i-built-an-hourly-newspaper-for-e-ink-and-turned-the-pipeline-into-an-mcp-server-1ebg","created_at":"2026-09-13T01:36:17.352356","last_synchronized":"2026-09-13T01:36:17.352356","sentiment":{"sentiment":"Positive","score":0.8876,"details":{"neg":0.029,"neu":0.924,"pos":0.047,"compound":0.8876}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"Set a Pause to Over-Engineering: Why AWS Lambda MicroVMs Will Serve Indie Developers in 2026","link":"https://dev.to/safiya_k/set-a-pause-to-over-engineering-why-aws-lambda-microvms-will-serve-indie-developers-in-2026-58nh","description":"<p>Introduction<br />\nFor a moment, let's be completely sincere with ourselves.<br />\nIf you read the latest enterprise tech blogs right now, you might think that every single person in the software industry is managing a huge fleet of hundreds of autonomous AI agents. Numerous announcements on multi-account registries, company governance dashboards, complicated tool policies, and prolonged orchestration runtimes have become increasingly common in the cloud world.<br />\nIt all sounds incredibly impressive if you are a corporate technology executive running a global team.<br />\nBut what if you are just a solo developer working on a side project?What if you are an indie hacker trying to build a fast MVP, or a startup engineer who simply wants to run a clean background job without setting up a massive compliance framework?<br />\nFor the everyday developer, the recent trend towards heavy enterprise infrastructure can feel exhausting.<br />\nThe good news is that AWS has not forgotten about the builders who just want to keep things small, fast, and simple.While one side of the cloud world went deep into complex corporate AI orchestration, the serverless team quietly shipped something that completely changes the game for indie developers: AWS Lambda MicroVMs.<br />\nThis new computing shift is the exact opposite of over-engineered enterprise software.<br />\nIt takes us back to what made the cloud fun in the first place: writing simple code, deploying it instantly, and letting the infrastructure handle the rest without charging you a fortune.</p>\n\n<p>Deviating from the Enterprise Over-Engineering Trap<br />\nOver the past couple of years, it feels like the software industry has accidentally re-invented the very infrastructure headaches we tried to escape.<br />\nDevelopers who wanted to build simple applications found themselves forced into adopting massive distributed microservices, heavy Kubernetes clusters, and multi-layered security frameworks before they even wrote their first user feature.<br />\nWhen you are a single-handed builder or part of a tiny team, your absolute highest priority is speed.<br />\nYou do not have the time to sit in compliance meetings, manage massive cross-account networks, or design complex system policies. You want to write a basic function, hook it up to an endpoint, and watch it work.<br />\nTraditional serverless computing with AWS Lambda was always supposed to deliver on this promise.<br />\nYou write your code, and AWS spins up a tiny container to execute it whenever a user hits your app. Nevertheless, traditional serverless has always had two major pain points that caused developers to start over-engineering their setups: cold starts and the total loss of temporary data storage between runs.<br />\nTo fix those two small issues, developers started spinning up full virtual servers and database instances.<br />\nSuddenly, a simple project turned into a complex architecture that cost real money every month just to keep running.</p>\n\n<p>Enter the MicroVM: Real Servers, Serverless Simplicity<br />\nThis is exactly why the release of AWS Lambda MicroVMs is such a massive deal for regular developers in 2026.<br />\nA MicroVM is an incredibly lightweight, completely isolated virtual machine sandbox.It gives you the raw power and security isolation of an individual dedicated server, but it retains the pure, effortless deployment model of a standard serverless function.<br />\nInstead of your code running in a shared container environment that instantly disappears the second your function finishes, a MicroVM creates a dedicated space just for your application session.<br />\nIt spins up in a matter of milliseconds, meaning cold starts are effectively dead, and it can preserve its internal state and temporary local files for up to eight continuous hours.<br />\nThis closes the massive structural gap that used to exist between basic serverless functions and heavy virtual machines.<br />\nYou no longer have to choose between a lightweight function that forgets everything instantly and an expensive, always-on cloud server that you have to manually patch and secure. You get the best of both worlds.<br />\nWhy This Architecture Shifts the Power Back to Solo Builders<br />\nThe appeal of this new framework is rooted in its straightforward design.<br />\nIt completely reverses the common enterprise approach that involves adding more layers, more dashboards, and more configuration options to your cloud environment. Here’s how it transforms the development experience for small teams and solo creators:</p>\n\n<p>1.True State Retention Without Complicated Databases<br />\nEarlier, if you wanted your backend code to retain a simple piece of information from a past user action, you had to store that data in an external database, such as Amazon DynamoDB.<br />\nFor a small application, setting up a database just to pass along a single temporary variable or session token felt like extra effort and complexity.</p>\n\n<p>With MicroVMs, the sandbox environment can maintain its memory state for up to eight hours, allowing you to store temporary data directly in the function’s memory cache.<br />\nIf a user is completing a multi-step task that spans an hour, your function will remember exactly where they left off, without needing to make any slow or costly database requests.This simplifies development, removes the need for complicated code loops, and reduces your cloud expenses.</p>\n\n<p>2.Total Isolation Without Complex Security Networks<br />\nIn large organizations, security teams often spend weeks configuring firewalls, identity tokens, and access policies to prevent applications from sharing or exposing data.<br />\nFor a solo developer, building such advanced security systems is a major roadblock.</p>\n\n<p>MicroVMs address this by offering complete isolation at the hardware level.<br />\nSince there is no shared kernel or memory between different user sessions, your application is inherently protected by default.You don’t need to spend hours configuring advanced security groups or go through lengthy compliance documents to safeguard your project from common web vulnerabilities.</p>\n\n<p>3.Rapid Launch Speeds and Zero Management<br />\nThe biggest challenge of using a traditional virtual cloud server is the need for ongoing maintenance.<br />\nYou must keep track of operating system updates, handle storage volumes, and pay a fixed fee every hour the server is online, even if no one is accessing your website.</p>\n\n<p>MicroVMs require no infrastructure management at all.<br />\nWhen a user accesses your application, the machine starts up instantly.Once the traffic stops, the system automatically shuts down completely.You only pay for the exact millisecond your code is running, yet you still enjoy the performance of a dedicated machine.</p>\n\n<p>Keeping Development Light, Fun, and Cost-Efficient<br />\nThe real advantage for developers lies in how this affects the financial side of project development.<br />\nEnterprise systems are built to support millions of users and require a large budget just to stay operational. However, when starting a new project, spending hundreds of dollars a month on cloud infrastructure before you have your first paying customer can quickly derail your efforts.</p>\n\n<p>By using ultra-lightweight serverless components, you can run an entire production-grade web application for very little cost.<br />\nThe mix of scale-to-zero compute pricing, minimal database use, and faster execution speeds means your regular operational costs are nearly zero.</p>\n\n<p>This financial flexibility changes how you approach development.<br />\nWhen there is no cost to running an application in the cloud, you are free to experiment, try out unusual side projects, and test new ideas without worrying about a large monthly cloud bill.<br />\nA Simple Roadmap for the Modern Solo Developer<br />\nIf you aim to move away from the over-engineering practices common in large enterprises and adopt a streamlined, efficient development process this year, your approach is quite clear:</p>\n\n<ul>\n<li>Focus on Code, Not Runtimes: Develop the core functions of your application as simple, readable code.\nLet the cloud infrastructure manage the underlying details of isolation and resources automatically.</li>\n<li><p>Use In-Memory Performance: Take advantage of the ability to store temporary user data locally for extended periods, rather than quickly building large and complex database systems.</p></li>\n<li><p>Design for Scale-to-Zero: Make sure your architecture is built around event-triggered processes.<br />\nIf your application isn’t being used at a certain time, your cloud costs should be minimal.</p></li>\n<li><p>Keep the Architecture Simple: Avoid adding extra layers such as orchestration services, dashboards, or complicated policy systems unless your application truly needs them.</p></li>\n</ul>\n\n<p>Conclusion:<br />\nThe Power of Less Software development has traditionally arrived in cycles.<br />\nWe often move from simple tools to overly complex frameworks, only to realize that this complexity is slowing us down and pushing us to return to simpler solutions.<br />\nWhile the enterprise world continues to create massive, highly managed systems that require entire teams to operate, the most successful independent developers are choosing a different path.<br />\nThey are building with minimal layers, using lightweight technologies like AWS Lambda MicroVMs to create faster, more cost-effective, and more intelligent applications.<br />\nUltimately, your users don’t care about the complexity of your cloud setup or whether you have advanced compliance tools running in the background.<br />\nThey care about your application being fast, dependable, and doing exactly what it promises. By choosing simplicity over complexity, you gain a significant advantage as a developer: the ability to create and deliver great ideas before your competitors even finish setting up their infrastructure.</p>","content":"Introduction\nFor a moment, let's be completely sincere with ourselves.\nIf you read the latest enterprise tech blogs right now, you might think that every single person in the software industry is managing a huge fleet of hundreds of autonomous AI agents. Numerous announcements on multi-account registries, company governance dashboards, complicated tool policies, and prolonged orchestration runtimes have become increasingly common in the cloud world.\nIt all sounds incredibly impressive if you are a corporate technology executive running a global team.\nBut what if you are just a solo developer working on a side project?What if you are an indie hacker trying to build a fast MVP, or a startup engineer who simply wants to run a clean background job without setting up a massive compliance framework?\nFor the everyday developer, the recent trend towards heavy enterprise infrastructure can feel exhausting.\nThe good news is that AWS has not forgotten about the builders who just want to keep things small, fast, and simple.While one side of the cloud world went deep into complex corporate AI orchestration, the serverless team quietly shipped something that completely changes the game for indie developers: AWS Lambda MicroVMs.\nThis new computing shift is the exact opposite of over-engineered enterprise software.\nIt takes us back to what made the cloud fun in the first place: writing simple code, deploying it instantly, and letting the infrastructure handle the rest without charging you a fortune.\nDeviating from the Enterprise Over-Engineering Trap\nOver the past couple of years, it feels like the software industry has accidentally re-invented the very infrastructure headaches we tried to escape.\nDevelopers who wanted to build simple applications found themselves forced into adopting massive distributed microservices, heavy Kubernetes clusters, and multi-layered security frameworks before they even wrote their first user feature.\nWhen you are a single-handed builder or part of a tiny team, your absolute highest priority is speed.\nYou do not have the time to sit in compliance meetings, manage massive cross-account networks, or design complex system policies. You want to write a basic function, hook it up to an endpoint, and watch it work.\nTraditional serverless computing with AWS Lambda was always supposed to deliver on this promise.\nYou write your code, and AWS spins up a tiny container to execute it whenever a user hits your app. Nevertheless, traditional serverless has always had two major pain points that caused developers to start over-engineering their setups: cold starts and the total loss of temporary data storage between runs.\nTo fix those two small issues, developers started spinning up full virtual servers and database instances.\nSuddenly, a simple project turned into a complex architecture that cost real money every month just to keep running.\nEnter the MicroVM: Real Servers, Serverless Simplicity\nThis is exactly why the release of AWS Lambda MicroVMs is such a massive deal for regular developers in 2026.\nA MicroVM is an incredibly lightweight, completely isolated virtual machine sandbox.It gives you the raw power and security isolation of an individual dedicated server, but it retains the pure, effortless deployment model of a standard serverless function.\nInstead of your code running in a shared container environment that instantly disappears the second your function finishes, a MicroVM creates a dedicated space just for your application session.\nIt spins up in a matter of milliseconds, meaning cold starts are effectively dead, and it can preserve its internal state and temporary local files for up to eight continuous hours.\nThis closes the massive structural gap that used to exist between basic serverless functions and heavy virtual machines.\nYou no longer have to choose between a lightweight function that forgets everything instantly and an expensive, always-on cloud server that you have to manually patch and secure. You get the best of both worlds.\nWhy This Architecture Shifts the Power Back to Solo Builders\nThe appeal of this new framework is rooted in its straightforward design.\nIt completely reverses the common enterprise approach that involves adding more layers, more dashboards, and more configuration options to your cloud environment. Here’s how it transforms the development experience for small teams and solo creators:\n1.True State Retention Without Complicated Databases\nEarlier, if you wanted your backend code to retain a simple piece of information from a past user action, you had to store that data in an external database, such as Amazon DynamoDB.\nFor a small application, setting up a database just to pass along a single temporary variable or session token felt like extra effort and complexity.\nWith MicroVMs, the sandbox environment can maintain its memory state for up to eight hours, allowing you to store temporary data directly in the function’s memory cache.\nIf a user is completing a multi-step task that spans an hour, your function will remember exactly where they left off, without needing to make any slow or costly database requests.This simplifies development, removes the need for complicated code loops, and reduces your cloud expenses.\n2.Total Isolation Without Complex Security Networks\nIn large organizations, security teams often spend weeks configuring firewalls, identity tokens, and access policies to prevent applications from sharing or exposing data.\nFor a solo developer, building such advanced security systems is a major roadblock.\nMicroVMs address this by offering complete isolation at the hardware level.\nSince there is no shared kernel or memory between different user sessions, your application is inherently protected by default.You don’t need to spend hours configuring advanced security groups or go through lengthy compliance documents to safeguard your project from common web vulnerabilities.\n3.Rapid Launch Speeds and Zero Management\nThe biggest challenge of using a traditional virtual cloud server is the need for ongoing maintenance.\nYou must keep track of operating system updates, handle storage volumes, and pay a fixed fee every hour the server is online, even if no one is accessing your website.\nMicroVMs require no infrastructure management at all.\nWhen a user accesses your application, the machine starts up instantly.Once the traffic stops, the system automatically shuts down completely.You only pay for the exact millisecond your code is running, yet you still enjoy the performance of a dedicated machine.\nKeeping Development Light, Fun, and Cost-Efficient\nThe real advantage for developers lies in how this affects the financial side of project development.\nEnterprise systems are built to support millions of users and require a large budget just to stay operational. However, when starting a new project, spending hundreds of dollars a month on cloud infrastructure before you have your first paying customer can quickly derail your efforts.\nBy using ultra-lightweight serverless components, you can run an entire production-grade web application for very little cost.\nThe mix of scale-to-zero compute pricing, minimal database use, and faster execution speeds means your regular operational costs are nearly zero.\nThis financial flexibility changes how you approach development.\nWhen there is no cost to running an application in the cloud, you are free to experiment, try out unusual side projects, and test new ideas without worrying about a large monthly cloud bill.\nA Simple Roadmap for the Modern Solo Developer\nIf you aim to move away from the over-engineering practices common in large enterprises and adopt a streamlined, efficient development process this year, your approach is quite clear:\n- Focus on Code, Not Runtimes: Develop the core functions of your application as simple, readable code. Let the cloud infrastructure manage the underlying details of isolation and resources automatically.\nUse In-Memory Performance: Take advantage of the ability to store temporary user data locally for extended periods, rather than quickly building large and complex database systems.\nDesign for Scale-to-Zero: Make sure your architecture is built around event-triggered processes.\nIf your application isn’t being used at a certain time, your cloud costs should be minimal.Keep the Architecture Simple: Avoid adding extra layers such as orchestration services, dashboards, or complicated policy systems unless your application truly needs them.\nConclusion:\nThe Power of Less Software development has traditionally arrived in cycles.\nWe often move from simple tools to overly complex frameworks, only to realize that this complexity is slowing us down and pushing us to return to simpler solutions.\nWhile the enterprise world continues to create massive, highly managed systems that require entire teams to operate, the most successful independent developers are choosing a different path.\nThey are building with minimal layers, using lightweight technologies like AWS Lambda MicroVMs to create faster, more cost-effective, and more intelligent applications.\nUltimately, your users don’t care about the complexity of your cloud setup or whether you have advanced compliance tools running in the background.\nThey care about your application being fast, dependable, and doing exactly what it promises. By choosing simplicity over complexity, you gain a significant advantage as a developer: the ability to create and deliver great ideas before your competitors even finish setting up their infrastructure.\nTop comments (0)","published":"Sat, 12 Sep 2026 22:47:01 +0000","author":"Safiya","guid":"https://dev.to/safiya_k/set-a-pause-to-over-engineering-why-aws-lambda-microvms-will-serve-indie-developers-in-2026-58nh","created_at":"2026-09-13T01:36:17.352356","last_synchronized":"2026-09-13T01:36:17.352356","sentiment":{"sentiment":"Positive","score":0.9989,"details":{"neg":0.046,"neu":0.838,"pos":0.116,"compound":0.9989}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"What has working with multiple languages been like on your projects?","link":"https://dev.to/n1x3n/what-has-working-with-multiple-languages-been-like-on-your-projects-2118","description":"<p>Hey everyone. I’m exploring a programming-language project and want to understand how people currently work across languages.</p>\n\n<p>Thinking about your most recent project where code in different languages needed to work together:</p>\n\n<ul>\n<li>What were you building, which languages were involved, and how did they communicate?</li>\n<li>Did anything go wrong at that connection? If so, what happened, how did you investigate it, and roughly how long did it take?</li>\n<li>What tools or approaches did you use, and how well did they work?</li>\n</ul>\n\n<p>If it was straightforward, I’d like to hear what made it work well too. One concrete example is enough.</p>","content":"Hey everyone. I’m exploring a programming-language project and want to understand how people currently work across languages.\nThinking about your most recent project where code in different languages needed to work together:\n- What were you building, which languages were involved, and how did they communicate?\n- Did anything go wrong at that connection? If so, what happened, how did you investigate it, and roughly how long did it take?\n- What tools or approaches did you use, and how well did they work?\nIf it was straightforward, I’d like to hear what made it work well too. One concrete example is enough.\nTop comments (0)","published":"Sat, 12 Sep 2026 22:44:34 +0000","author":"Aliameen Adebukola Fatunbi","guid":"https://dev.to/n1x3n/what-has-working-with-multiple-languages-been-like-on-your-projects-2118","created_at":"2026-09-13T01:36:17.352356","last_synchronized":"2026-09-13T01:36:17.352356","sentiment":{"sentiment":"Positive","score":0.594,"details":{"neg":0.028,"neu":0.889,"pos":0.082,"compound":0.594}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"Oracle Integration Cloud Integration Patterns","link":"https://dev.to/someshp/oracle-integration-cloud-integration-patterns-fnc","description":"<p>Twelve reusable patterns, real-world examples, and implementation guidance for Oracle Integration 3</p>\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv6v5ly87c4zhc419nz52.png\"><img alt=\" \" height=\"98\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv6v5ly87c4zhc419nz52.png\" width=\"800\" /></a><br />\n<strong>About this field guide</strong> </p>\n\n<p>I wrote this guide to connect Oracle Integration features with the architecture decisions we make on real projects. It presents twelve practical patterns, each supported by an enterprise or healthcare example, an architecture view and a sequence diagram. The aim is to make the patterns easy to recognize, compare and apply - not simply describe product features. </p>\n\n<p><strong>How to read this guide</strong> </p>\n\n<p>These patterns are building blocks rather than isolated templates. A <br />\nproduction solution may combine an API facade, orchestration, asynchronous hand-off, and a parking lot. Start with the business interaction, then use the diagrams to understand how OIC and the participating systems work together. </p>\n\n<ol>\n<li>*<em>Synchronous request–reply Pattern *</em>\nUse when the caller needs an immediate business response. A REST/SOAP trigger validates and maps the request, invokes one or more applications, and returns a consolidated response. Keep the flow short because the client connection remains open. </li>\n</ol>\n\n<p>Example. A provider portal submits an eligibility inquiry. OIC calls a payer eligibility API and returns coverage, copay and deductible in the same interaction. </p>\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyzgzi81wzf1kxue3xukx.png\"><img alt=\" \" height=\"255\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyzgzi81wzf1kxue3xukx.png\" width=\"800\" /></a></p>\n\n<ol>\n<li>*<em>API facade and protocol mediation Pattern *</em>\nPlace OIC between consumers and a legacy or SaaS interface to hide protocol, schema, and authentication differences. The facade exposes a stable contract while mappings and adapters absorb downstream change. </li>\n</ol>\n\n<p>Example. A mobile claims app sends JSON/REST. OIC transforms it to the SOAP/XML contract required by a legacy claims platform, then converts the response back to JSON. </p>\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fguxv0fcsl3uar79dsurw.png\"><img alt=\" \" height=\"255\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fguxv0fcsl3uar79dsurw.png\" width=\"800\" /></a></p>\n\n<ol>\n<li>*<em>Content-based routing Pattern *</em>\nInspect payload or header values and use switch branches to send each message to the correct target. Centralize routing rules and define a default path so unexpected values are observable rather than silently discarded. </li>\n</ol>\n\n<p>Example. A claim is routed to the commercial, Medicare, or Medicaid pricing service according to line of business and state. </p>\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffrhtzlyzm6klalflg50t.png\"><img alt=\" \" height=\"283\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffrhtzlyzm6klalflg50t.png\" width=\"799\" /></a></p>\n\n<p>4.<strong>Orchestration and aggregation Pattern</strong> <br />\nCoordinate a multi-step business transaction, including sequential invokes, enrichment, decisions, and response assembly. Use scopes for local fault handling; avoid holding a synchronous caller through long-running work. </p>\n\n<p>Example. Member onboarding creates the member in CRM, validates identity, enrolls benefits, and combines the generated identifiers into one completion record. </p>\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvueb3ujelv5duo9b74af.png\"><img alt=\" \" height=\"285\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvueb3ujelv5duo9b74af.png\" width=\"800\" /></a></p>\n\n<p>5.*<em>Parallel fan-out and gather Pattern *</em><br />\nInvoke independent services concurrently and aggregate their results. This reduces elapsed time when branches do not depend on one another. Define how partial failures are reported and whether every branch is mandatory. </p>\n\n<p>Example. Before claim adjudication, OIC obtains eligibility, provider status, and authorization in parallel, then builds a single validation result.</p>\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg2z9pfmhbotz1bhxvzpb.png\"><img alt=\" \" height=\"269\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg2z9pfmhbotz1bhxvzpb.png\" width=\"800\" /></a></p>\n\n<p>6.*<em>Asynchronous hand-off Pattern *</em><br />\nAcknowledge quickly, then perform slow or high-volume processing outside the original request. Use a queue, event service, or one-way invocation, and carry a correlation ID for status and callbacks. </p>\n\n<p>Example. A hospital submits a large encounter batch and receives HTTP 202 with a tracking ID. OIC processes records asynchronously and later posts completion status. </p>\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm3e59ff1okqevb5k1gac.png\"><img alt=\" \" height=\"260\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm3e59ff1okqevb5k1gac.png\" width=\"800\" /></a></p>\n\n<p>7.*<em>Publish–subscribe Pattern *</em><br />\nPublish one canonical business event and let multiple subscriber integrations react independently. This reduces point-to-point coupling and allows new consumers without changing the publisher. Design subscribers to be idempotent. </p>\n\n<p>Example. A Patient Update event independently updates CRM, analytics and care-management systems through Oracle Integration Messaging or an external event backbone.</p>\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp3j9011mnv59bmtn32t5.png\"><img alt=\" \" height=\"270\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp3j9011mnv59bmtn32t5.png\" width=\"800\" /></a></p>\n\n<p>8.*<em>Event-driven integration Pattern *</em><br />\nTrigger processing when a business or cloud event occurs instead of polling. Events may originate in SaaS applications, OCI Events, streaming, or messaging services. Include filtering, deduplication, and replay strategy. </p>\n\n<p>Example. An object-created event for a remittance file starts validation and posting immediately after the file lands in OCI Object Storage. </p>\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz116zr432kdwn3sf64hl.png\"><img alt=\" \" height=\"238\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz116zr432kdwn3sf64hl.png\" width=\"799\" /></a></p>\n\n<p>9.*<em>Scheduled polling and incremental synchronization Pattern *</em><br />\nRun on a defined schedule when the source cannot emit events. Query records changed since a stored high-water mark, process bounded pages, and advanced the watermark only after successful completion. </p>\n\n<p>Example. Every 15 minutes OIC reads newly changed provider records from an on-premises database through the connectivity agent and upserts them into Oracle Fusion Cloud. </p>\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxt08vtu9xiopdii6ezwe.png\"><img alt=\" \" height=\"259\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxt08vtu9xiopdii6ezwe.png\" width=\"799\" /></a></p>\n\n<p>10.*<em>Bulk data and file transfer Pattern *</em><br />\nMove large datasets through files rather than chatty record-level calls. OIC can use FTP/SFTP, File Adapter, Object Storage and Stage File actions to list, read, write, zip, unzip, encrypt or decrypt content. Stream or segment large files. </p>\n\n<p>Example. A nightly enrollment CSV is collected from SFTP, decrypted, split into manageable batches, transformed and loaded into the benefits platform. Rejected records go to a separate file.</p>\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzaxgh8tcrgelt7pditl2.png\"><img alt=\" \" height=\"263\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzaxgh8tcrgelt7pditl2.png\" width=\"799\" /></a></p>\n\n<ol>\n<li>*<em>B2B / EDI gateway Pattern *</em>\nUse trading partner agreements, document definitions, and transport settings to exchange X12 or EDIFACT. Translate between EDI and canonical application messages, validate envelopes, track acknowledgements, and retain business-level visibility. </li>\n</ol>\n\n<p>Example. A health plan receives X12 837 claims over AS2, validates and translates them, submits canonical claims internally, and returns the appropriate acknowledgement. </p>\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8cxfovorcyfza911s5u9.png\"><img alt=\" \" height=\"268\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8cxfovorcyfza911s5u9.png\" width=\"800\" /></a></p>\n\n<ol>\n<li>*<em>Reliable delivery, retry and parking lot Pattern *</em>\nWrap volatile invokes in scopes, classify faults and retry only transient failures with controlled backoff. Persist exhausted or data-related failures in a parking-lot store with payload, error and correlation data for correction and replay. </li>\n</ol>\n\n<p>Example. If a provider API times out, OIC retries. After the limit, it stores the request in ATP, alerts operations and allows for safe resubmission after correction. </p>\n\n<p><a class=\"article-body-image-wrapper\" href=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqr5cmmbpdsov86rqhuwy.png\"><img alt=\" \" height=\"253\" src=\"https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqr5cmmbpdsov86rqhuwy.png\" width=\"800\" /></a></p>\n\n<p>*<em>Choosing and combining patterns *</em></p>\n\n<p>Start with the interaction contract.</p>\n\n<ul>\n<li><p>Use synchronous request–reply only when the answer is fast and required immediately. </p></li>\n<li><p>Use asynchronous hand-off for long running or bursty work.</p></li>\n<li><p>Prefer events or publish–subscribe when several consumers should evolve independently.</p></li>\n<li><p>Use scheduling only when the source cannot notify OIC and protect scheduled extraction with watermarks and paging.</p></li>\n<li><p>Use files for genuine bulk payloads and B2B when partner agreements, EDI validation, and acknowledgements matter.</p></li>\n<li><p>Across every choice, design idempotency, correlation, security, observability, fault classification and replay from the beginning. </p></li>\n</ul>\n\n<p>*<em>Architecture guardrails *</em></p>\n\n<ul>\n<li><p>Keep canonical messages business focused. Do not expose one application’s schema as the enterprise contract. </p></li>\n<li><p>Separate real-time APIs from long-running work, and do not chain many synchronous dependencies. </p></li>\n<li><p>Track a business identifier and correlation ID across every hop and never rely only on an OIC instance ID. </p></li>\n<li><p>Make consumers and replay operations idempotent, especially for events, retries and file reprocessing. </p></li>\n<li><p>Use the connectivity agent for private endpoints and keep secrets in managed connection/security policies. </p></li>\n<li><p>Define operational ownership, alert thresholds, retry limits, parking-lot triage, retention, and recovery objectives. </p></li>\n</ul>\n\n<p>*<em>Oracle references *</em></p>\n\n<ul>\n<li><p>Understand Integration Styles </p></li>\n<li><p>Publish and Subscribe with Oracle Integration Messaging</p></li>\n<li><p>Create Scheduled Integrations</p></li>\n<li><p>Stage File Processing</p></li>\n<li><p>Parallel Processing</p></li>\n<li><p>OIC Design Best Practices</p></li>\n</ul>","content":"Twelve reusable patterns, real-world examples, and implementation guidance for Oracle Integration 3\nI wrote this guide to connect Oracle Integration features with the architecture decisions we make on real projects. It presents twelve practical patterns, each supported by an enterprise or healthcare example, an architecture view and a sequence diagram. The aim is to make the patterns easy to recognize, compare and apply - not simply describe product features.\nHow to read this guide\nThese patterns are building blocks rather than isolated templates. A\nproduction solution may combine an API facade, orchestration, asynchronous hand-off, and a parking lot. Start with the business interaction, then use the diagrams to understand how OIC and the participating systems work together.\n- *Synchronous request–reply Pattern * Use when the caller needs an immediate business response. A REST/SOAP trigger validates and maps the request, invokes one or more applications, and returns a consolidated response. Keep the flow short because the client connection remains open.\nExample. A provider portal submits an eligibility inquiry. OIC calls a payer eligibility API and returns coverage, copay and deductible in the same interaction.\n- *API facade and protocol mediation Pattern * Place OIC between consumers and a legacy or SaaS interface to hide protocol, schema, and authentication differences. The facade exposes a stable contract while mappings and adapters absorb downstream change.\nExample. A mobile claims app sends JSON/REST. OIC transforms it to the SOAP/XML contract required by a legacy claims platform, then converts the response back to JSON.\n- *Content-based routing Pattern * Inspect payload or header values and use switch branches to send each message to the correct target. Centralize routing rules and define a default path so unexpected values are observable rather than silently discarded.\nExample. A claim is routed to the commercial, Medicare, or Medicaid pricing service according to line of business and state.\n4.Orchestration and aggregation Pattern\nCoordinate a multi-step business transaction, including sequential invokes, enrichment, decisions, and response assembly. Use scopes for local fault handling; avoid holding a synchronous caller through long-running work.\nExample. Member onboarding creates the member in CRM, validates identity, enrolls benefits, and combines the generated identifiers into one completion record.\n5.*Parallel fan-out and gather Pattern *\nInvoke independent services concurrently and aggregate their results. This reduces elapsed time when branches do not depend on one another. Define how partial failures are reported and whether every branch is mandatory.\nExample. Before claim adjudication, OIC obtains eligibility, provider status, and authorization in parallel, then builds a single validation result.\n6.*Asynchronous hand-off Pattern *\nAcknowledge quickly, then perform slow or high-volume processing outside the original request. Use a queue, event service, or one-way invocation, and carry a correlation ID for status and callbacks.\nExample. A hospital submits a large encounter batch and receives HTTP 202 with a tracking ID. OIC processes records asynchronously and later posts completion status.\n7.*Publish–subscribe Pattern *\nPublish one canonical business event and let multiple subscriber integrations react independently. This reduces point-to-point coupling and allows new consumers without changing the publisher. Design subscribers to be idempotent.\nExample. A Patient Update event independently updates CRM, analytics and care-management systems through Oracle Integration Messaging or an external event backbone.\n8.*Event-driven integration Pattern *\nTrigger processing when a business or cloud event occurs instead of polling. Events may originate in SaaS applications, OCI Events, streaming, or messaging services. Include filtering, deduplication, and replay strategy.\nExample. An object-created event for a remittance file starts validation and posting immediately after the file lands in OCI Object Storage.\n9.*Scheduled polling and incremental synchronization Pattern *\nRun on a defined schedule when the source cannot emit events. Query records changed since a stored high-water mark, process bounded pages, and advanced the watermark only after successful completion.\nExample. Every 15 minutes OIC reads newly changed provider records from an on-premises database through the connectivity agent and upserts them into Oracle Fusion Cloud.\n10.*Bulk data and file transfer Pattern *\nMove large datasets through files rather than chatty record-level calls. OIC can use FTP/SFTP, File Adapter, Object Storage and Stage File actions to list, read, write, zip, unzip, encrypt or decrypt content. Stream or segment large files.\nExample. A nightly enrollment CSV is collected from SFTP, decrypted, split into manageable batches, transformed and loaded into the benefits platform. Rejected records go to a separate file.\n- *B2B / EDI gateway Pattern * Use trading partner agreements, document definitions, and transport settings to exchange X12 or EDIFACT. Translate between EDI and canonical application messages, validate envelopes, track acknowledgements, and retain business-level visibility.\nExample. A health plan receives X12 837 claims over AS2, validates and translates them, submits canonical claims internally, and returns the appropriate acknowledgement.\n- *Reliable delivery, retry and parking lot Pattern * Wrap volatile invokes in scopes, classify faults and retry only transient failures with controlled backoff. Persist exhausted or data-related failures in a parking-lot store with payload, error and correlation data for correction and replay.\nExample. If a provider API times out, OIC retries. After the limit, it stores the request in ATP, alerts operations and allows for safe resubmission after correction.\n*Choosing and combining patterns *\nStart with the interaction contract.\nUse synchronous request–reply only when the answer is fast and required immediately.\nUse asynchronous hand-off for long running or bursty work.\nPrefer events or publish–subscribe when several consumers should evolve independently.\nUse scheduling only when the source cannot notify OIC and protect scheduled extraction with watermarks and paging.\nUse files for genuine bulk payloads and B2B when partner agreements, EDI validation, and acknowledgements matter.\nAcross every choice, design idempotency, correlation, security, observability, fault classification and replay from the beginning.\n*Architecture guardrails *\nKeep canonical messages business focused. Do not expose one application’s schema as the enterprise contract.\nSeparate real-time APIs from long-running work, and do not chain many synchronous dependencies.\nTrack a business identifier and correlation ID across every hop and never rely only on an OIC instance ID.\nMake consumers and replay operations idempotent, especially for events, retries and file reprocessing.\nUse the connectivity agent for private endpoints and keep secrets in managed connection/security policies.\nDefine operational ownership, alert thresholds, retry limits, parking-lot triage, retention, and recovery objectives.\n*Oracle references *\nUnderstand Integration Styles\nPublish and Subscribe with Oracle Integration Messaging\nCreate Scheduled Integrations\nStage File Processing\nParallel Processing\nOIC Design Best Practices\nTop comments (0)","published":"Sat, 12 Sep 2026 22:42:52 +0000","author":"Somesh Purohit","guid":"https://dev.to/someshp/oracle-integration-cloud-integration-patterns-fnc","created_at":"2026-09-13T01:36:17.352356","last_synchronized":"2026-09-13T01:36:17.352356","sentiment":{"sentiment":"Positive","score":0.9695,"details":{"neg":0.03,"neu":0.921,"pos":0.05,"compound":0.9695}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"Architecting a Low-Power Geofencing Engine for Android","link":"https://dev.to/haseebthedev0/architecting-a-low-power-geofencing-engine-for-android-1e6j","description":"<p>It happened during a quiet, mid-afternoon meeting. The room was deathly silent, the air thick with the weight of a quarterly review, when my phone decided to belt out an aggressive, high-decibel ringtone. My face turned crimson. I scrambled to silence it, fumbling with the volume rockers, but the damage was done. The rhythm of the meeting was shattered, and I spent the next ten minutes apologizing rather than contributing. That was the moment I realized my phone, for all its intelligence, was failing me at the most basic level of etiquette.</p>\n\n<p>We live in a world of constant digital noise, yet our devices lack the context-awareness to know when to shut up. I found myself manually toggling between vibrate, silent, and normal modes dozens of times a day. If I forgot to unmute after a gym session or a movie, I’d miss critical calls. If I forgot to silence before a lecture, I’d be the source of distraction. Existing solutions often felt bloated, requiring cloud syncs or constant battery-draining polling that made my phone feel sluggish. I didn't want a suite of features I’d never use; I just wanted my phone to know where I was and act accordingly without me needing to touch it.</p>\n\n<p>I sat down to build a tool that could handle this reliably. The core requirement was clear: it needed to be fully offline, privacy-focused, and battery-efficient. I realized that a simple time-based scheduler wasn't enough. Many of us operate on location-based habits—the gym, the office, the library. This led me to implement a geofencing engine. My primary concern was the trade-off between location accuracy and battery longevity. Continuous GPS tracking is an absolute battery killer, and I knew that if my users saw their battery drain by 20% in an afternoon, they would uninstall the app immediately.</p>\n\n<p>I opted for the <code>GeofencingClient</code> within the Google Play Services Location API. This approach is superior to manual location polling because it offloads the monitoring to the system. By defining circular regions (geofences), the system handles the heavy lifting of location updates, waking up the app only when a transition (entering or exiting) occurs. However, there is a catch: the accuracy of these geofences depends on the phone’s signal environment. In dense urban areas with tall buildings, GPS signal bouncing can cause 'false exits' where the system thinks you've left a building when you haven't. </p>\n\n<p>kotlin<br />\nval geofence = Geofence.Builder()<br />\n.setRequestId(id)<br />\n.setCircularRegion(lat, lon, radiusInMeters)<br />\n.setExpirationDuration(Geofence.NEVER_EXPIRE)<br />\n.setTransitionTypes(Geofence.GEOFENCE_TRANSITION_ENTER or Geofence.GEOFENCE_TRANSITION_EXIT)<br />\n.build()</p>\n\n<p>I had to implement a hysteresis buffer to combat this. Instead of reacting instantly to an exit event, I introduced a small timer that checks if the device remains outside the zone for more than 60 seconds. This simple architectural delay prevented the constant toggling of sound profiles when a user just moves to the other side of a large office building. I coupled this with an <code>IntentService</code> that handles the <code>GeofencingEvent</code>, ensuring the logic is processed in the background even if the main UI is closed. Keeping this entire stack local meant I had to manage state manually using <code>Room</code> for persistence, ensuring that after a reboot, the <code>AlarmManager</code> and <code>GeofencingClient</code> were correctly re-registered to restore the user's active sound routines.</p>\n\n<p>What surprised me most was the fragility of background execution on modern Android versions. My initial assumption was that if a user granted location permissions, my service would hum along indefinitely. I was wrong. Android’s 'Doze' mode and manufacturer-specific battery optimizations are aggressive. My early tests showed that on some devices, the geofencing triggers were delayed by up to twenty minutes because the OS prioritized saving power over my background listener. I learned that for critical routines, I couldn't rely solely on the system's geofencing triggers. I had to implement a fallback check.</p>\n\n<p>I eventually added a feature that triggers a short-lived foreground service upon a geofence event. By showing a notification, I effectively promoted my app from a 'background task' to a 'visible operation' in the eyes of the Android task manager, which drastically improved the reliability of sound profile changes. If I were starting over, I would have focused on the 'emergency bypass' feature much earlier. I initially thought silent mode should be absolute, but I realized that users are terrified of missing calls from family. Allowing specific contacts to override the silence was the single most requested feature in my early alpha tests. I also underestimated the complexity of time zones; if a user travels, their locally stored routine times can become completely misaligned. I had to shift my entire storage architecture to store UTC timestamps and calculate offsets locally based on the device's current locale.</p>\n\n<p>As you architect your own background systems, the biggest takeaway is to respect the user's battery as much as you respect their privacy. Don't build a 'polling-based' system if an 'event-based' system exists in the platform's APIs. The platform developers at Google put significant work into optimizing APIs like <code>GeofencingClient</code> for a reason; trying to roll your own location listener using <code>LocationManager</code> is almost always a mistake unless you have a hyper-specific use case that requires it. Always assume the system will kill your background process at the worst possible time, and design your state persistence so that your app can recover gracefully without the user needing to intervene.</p>\n\n<p>Testing on a wide range of devices—specifically cheaper, 'budget' Android phones—is non-negotiable. These devices often have the most aggressive background management policies, and if your code works there, it will work anywhere. Muffle was born out of my own frustration with these exact constraints, and it has evolved into a tool that keeps my phone silent when I need it to be, and audible when it matters. If you are interested in how I implemented the logic for prayer times alongside these geofences, you can find the project here: <a href=\"https://play.google.com/store/apps/details?id=com.muffle.app\" rel=\"noopener noreferrer\">https://play.google.com/store/apps/details?id=com.muffle.app</a></p>","content":"It happened during a quiet, mid-afternoon meeting. The room was deathly silent, the air thick with the weight of a quarterly review, when my phone decided to belt out an aggressive, high-decibel ringtone. My face turned crimson. I scrambled to silence it, fumbling with the volume rockers, but the damage was done. The rhythm of the meeting was shattered, and I spent the next ten minutes apologizing rather than contributing. That was the moment I realized my phone, for all its intelligence, was failing me at the most basic level of etiquette.\nWe live in a world of constant digital noise, yet our devices lack the context-awareness to know when to shut up. I found myself manually toggling between vibrate, silent, and normal modes dozens of times a day. If I forgot to unmute after a gym session or a movie, I’d miss critical calls. If I forgot to silence before a lecture, I’d be the source of distraction. Existing solutions often felt bloated, requiring cloud syncs or constant battery-draining polling that made my phone feel sluggish. I didn't want a suite of features I’d never use; I just wanted my phone to know where I was and act accordingly without me needing to touch it.\nI sat down to build a tool that could handle this reliably. The core requirement was clear: it needed to be fully offline, privacy-focused, and battery-efficient. I realized that a simple time-based scheduler wasn't enough. Many of us operate on location-based habits—the gym, the office, the library. This led me to implement a geofencing engine. My primary concern was the trade-off between location accuracy and battery longevity. Continuous GPS tracking is an absolute battery killer, and I knew that if my users saw their battery drain by 20% in an afternoon, they would uninstall the app immediately.\nI opted for the GeofencingClient\nwithin the Google Play Services Location API. This approach is superior to manual location polling because it offloads the monitoring to the system. By defining circular regions (geofences), the system handles the heavy lifting of location updates, waking up the app only when a transition (entering or exiting) occurs. However, there is a catch: the accuracy of these geofences depends on the phone’s signal environment. In dense urban areas with tall buildings, GPS signal bouncing can cause 'false exits' where the system thinks you've left a building when you haven't.\nkotlin\nval geofence = Geofence.Builder()\n.setRequestId(id)\n.setCircularRegion(lat, lon, radiusInMeters)\n.setExpirationDuration(Geofence.NEVER_EXPIRE)\n.setTransitionTypes(Geofence.GEOFENCE_TRANSITION_ENTER or Geofence.GEOFENCE_TRANSITION_EXIT)\n.build()\nI had to implement a hysteresis buffer to combat this. Instead of reacting instantly to an exit event, I introduced a small timer that checks if the device remains outside the zone for more than 60 seconds. This simple architectural delay prevented the constant toggling of sound profiles when a user just moves to the other side of a large office building. I coupled this with an IntentService\nthat handles the GeofencingEvent\n, ensuring the logic is processed in the background even if the main UI is closed. Keeping this entire stack local meant I had to manage state manually using Room\nfor persistence, ensuring that after a reboot, the AlarmManager\nand GeofencingClient\nwere correctly re-registered to restore the user's active sound routines.\nWhat surprised me most was the fragility of background execution on modern Android versions. My initial assumption was that if a user granted location permissions, my service would hum along indefinitely. I was wrong. Android’s 'Doze' mode and manufacturer-specific battery optimizations are aggressive. My early tests showed that on some devices, the geofencing triggers were delayed by up to twenty minutes because the OS prioritized saving power over my background listener. I learned that for critical routines, I couldn't rely solely on the system's geofencing triggers. I had to implement a fallback check.\nI eventually added a feature that triggers a short-lived foreground service upon a geofence event. By showing a notification, I effectively promoted my app from a 'background task' to a 'visible operation' in the eyes of the Android task manager, which drastically improved the reliability of sound profile changes. If I were starting over, I would have focused on the 'emergency bypass' feature much earlier. I initially thought silent mode should be absolute, but I realized that users are terrified of missing calls from family. Allowing specific contacts to override the silence was the single most requested feature in my early alpha tests. I also underestimated the complexity of time zones; if a user travels, their locally stored routine times can become completely misaligned. I had to shift my entire storage architecture to store UTC timestamps and calculate offsets locally based on the device's current locale.\nAs you architect your own background systems, the biggest takeaway is to respect the user's battery as much as you respect their privacy. Don't build a 'polling-based' system if an 'event-based' system exists in the platform's APIs. The platform developers at Google put significant work into optimizing APIs like GeofencingClient\nfor a reason; trying to roll your own location listener using LocationManager\nis almost always a mistake unless you have a hyper-specific use case that requires it. Always assume the system will kill your background process at the worst possible time, and design your state persistence so that your app can recover gracefully without the user needing to intervene.\nTesting on a wide range of devices—specifically cheaper, 'budget' Android phones—is non-negotiable. These devices often have the most aggressive background management policies, and if your code works there, it will work anywhere. Muffle was born out of my own frustration with these exact constraints, and it has evolved into a tool that keeps my phone silent when I need it to be, and audible when it matters. If you are interested in how I implemented the logic for prayer times alongside these geofences, you can find the project here: https://play.google.com/store/apps/details?id=com.muffle.app\nTop comments (0)","published":"Sat, 12 Sep 2026 22:35:21 +0000","author":"Haseeb","guid":"https://dev.to/haseebthedev0/architecting-a-low-power-geofencing-engine-for-android-1e6j","created_at":"2026-09-13T01:36:17.352356","last_synchronized":"2026-09-13T01:36:17.352356","sentiment":{"sentiment":"Negative","score":-0.938,"details":{"neg":0.084,"neu":0.843,"pos":0.072,"compound":-0.938}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"I Rewrote One Function in Squirrel, Then in C. One Line of Python Was Just as Fast.","link":"https://dev.to/natuworkguy/i-rewrote-one-function-in-squirrel-then-in-c-one-line-of-python-was-just-as-fast-3han","description":"<p>I maintain a 2D game engine in Python called ABS Engine. Last week I decided it needed C.</p>\n\n<p>Not because anything was slow. That is the embarrassing part. I wanted to drop a <code>.c</code> file into a folder and call it from Python with no build step anyone has to know about. It took 197 lines, most of them docstrings.</p>\n\n<p>The first function I moved over was <code>clamp</code>. Keep a number between a low and a high. Three comparisons.</p>\n\n<p>It had already been rewritten once. Two weeks earlier I had moved that same function out of Python and into Squirrel.</p>\n\n<p>So the real history of <code>clamp</code> in this engine is Python, then Squirrel, then C, for a function whose entire body is three comparisons.</p>\n\n<h2>\n\n\nThe Squirrel detour\n</h2>\n\n<p>The engine already embedded Tcl, because the editor is Tk and Tk is Tcl underneath. Squirrel came next, and <code>clamp</code> is what I used to find out whether an embedded VM was practical. Here is <code>engine/nut/math.nut</code>, in full:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight javascript\"><code><span class=\"kd\">function</span> <span class=\"nf\">clamp</span><span class=\"p\">(</span><span class=\"nx\">value</span><span class=\"p\">,</span> <span class=\"nx\">low</span><span class=\"p\">,</span> <span class=\"nx\">high</span><span class=\"p\">)</span> <span class=\"p\">{</span>\n<span class=\"k\">if </span><span class=\"p\">(</span><span class=\"nx\">value</span> <span class=\"o\">&lt;</span> <span class=\"nx\">low</span><span class=\"p\">)</span> <span class=\"p\">{</span>\n<span class=\"k\">return</span> <span class=\"nx\">low</span>\n<span class=\"p\">}</span>\n\n<span class=\"k\">if </span><span class=\"p\">(</span><span class=\"nx\">value</span> <span class=\"o\">&gt;</span> <span class=\"nx\">high</span><span class=\"p\">)</span> <span class=\"p\">{</span>\n<span class=\"k\">return</span> <span class=\"nx\">high</span>\n<span class=\"p\">}</span>\n\n<span class=\"k\">return</span> <span class=\"nx\">value</span>\n<span class=\"p\">}</span>\n</code></pre>\n\n</div>\n\n\n\n<p>And here is how Python called it:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight python\"><code><span class=\"k\">return</span> <span class=\"nf\">float</span><span class=\"p\">(</span><span class=\"nf\">nut_call_function</span><span class=\"p\">(</span><span class=\"sh\">\"</span><span class=\"s\">clamp</span><span class=\"sh\">\"</span><span class=\"p\">,</span> <span class=\"nf\">float</span><span class=\"p\">(</span><span class=\"n\">value</span><span class=\"p\">),</span> <span class=\"nf\">float</span><span class=\"p\">(</span><span class=\"n\">low</span><span class=\"p\">),</span> <span class=\"nf\">float</span><span class=\"p\">(</span><span class=\"n\">high</span><span class=\"p\">)))</span>\n</code></pre>\n\n</div>\n\n\n\n<p>Read that line twice, because it gives the whole game away. Three <code>float()</code> calls on the way in. A lookup in the Squirrel root table. A trip through an interpreter loop to run three comparisons. One more <code>float()</code> coming back. The docstring above it, and I am quoting my own repository, proudly said <code>*Implemented in Squirrel*</code>.</p>\n\n<h2>\n\n\nThe header is the contract\n</h2>\n\n<p>The rule I set for C was that adding a function should cost two files and zero configuration. Write <code>geometry.c</code>, write <code>geometry.h</code> next to it, call it. No setup.py entry, no CMake, no remembering to recompile.</p>\n\n<p>Here is <code>mathutil.h</code>, all of it:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight c\"><code><span class=\"kt\">double</span> <span class=\"nf\">clamp</span><span class=\"p\">(</span><span class=\"kt\">double</span> <span class=\"n\">value</span><span class=\"p\">,</span> <span class=\"kt\">double</span> <span class=\"n\">low</span><span class=\"p\">,</span> <span class=\"kt\">double</span> <span class=\"n\">high</span><span class=\"p\">);</span>\n</code></pre>\n\n</div>\n\n\n\n<p>The loader hands that text straight to cffi:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight python\"><code><span class=\"n\">ffibuilder</span><span class=\"p\">.</span><span class=\"nf\">cdef</span><span class=\"p\">(</span><span class=\"n\">header_path</span><span class=\"p\">.</span><span class=\"nf\">read_text</span><span class=\"p\">(</span><span class=\"n\">encoding</span><span class=\"o\">=</span><span class=\"sh\">\"</span><span class=\"s\">utf-8</span><span class=\"sh\">\"</span><span class=\"p\">))</span>\n<span class=\"n\">ffibuilder</span><span class=\"p\">.</span><span class=\"nf\">set_source</span><span class=\"p\">(</span>\n<span class=\"n\">module_name</span><span class=\"p\">,</span>\n<span class=\"n\">source_path</span><span class=\"p\">.</span><span class=\"nf\">read_text</span><span class=\"p\">(</span><span class=\"n\">encoding</span><span class=\"o\">=</span><span class=\"sh\">\"</span><span class=\"s\">utf-8</span><span class=\"sh\">\"</span><span class=\"p\">),</span>\n<span class=\"n\">include_dirs</span><span class=\"o\">=</span><span class=\"p\">[</span><span class=\"nf\">str</span><span class=\"p\">(</span><span class=\"n\">C_DIR</span><span class=\"p\">)],</span>\n<span class=\"n\">libraries</span><span class=\"o\">=</span><span class=\"p\">[]</span> <span class=\"k\">if</span> <span class=\"n\">sys</span><span class=\"p\">.</span><span class=\"n\">platform</span> <span class=\"o\">==</span> <span class=\"sh\">\"</span><span class=\"s\">win32</span><span class=\"sh\">\"</span> <span class=\"k\">else</span> <span class=\"p\">[</span><span class=\"sh\">\"</span><span class=\"s\">m</span><span class=\"sh\">\"</span><span class=\"p\">],</span>\n<span class=\"p\">)</span>\n</code></pre>\n\n</div>\n\n\n\n<p>That is the trick, and it is also the biggest gotcha in the project, so let me be loud about it. <code>cdef</code> is not a compiler and not a preprocessor. It parses a narrow subset of C declarations. An <code>#include</code> fails. Include guards fail, because <code>#ifndef</code> means nothing to it.</p>\n\n<p>So headers here are declarations and nothing else. The <code>.c</code> file includes its own header normally, because the real compiler handles that file and is fine with all of it. Two readers, two sets of rules. The header is the part Python is allowed to see.</p>\n\n<p>Rebuilds are decided by mtime: if a compiled module exists and is newer than both the <code>.c</code> and the <code>.h</code>, it gets imported, otherwise it gets rebuilt. <code>functools.cache</code> on top means you pay that check once per file per process. The practical effect is the thing I wanted. You edit the C, you hit Run, the new C is live. You never type the word \"build.\"</p>\n\n<p>One error message worth stealing. If the compile succeeds but the import fails, it almost always means the build folder holds a binary from a different Python, so the loader says exactly that and names the fix:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight python\"><code><span class=\"k\">raise</span> <span class=\"nc\">ModuleNotFoundError</span><span class=\"p\">(</span>\n<span class=\"sa\">f</span><span class=\"sh\">\"</span><span class=\"s\">Compiled </span><span class=\"si\">{</span><span class=\"n\">source_name</span><span class=\"si\">}</span><span class=\"s\">, but </span><span class=\"si\">{</span><span class=\"n\">module_name</span><span class=\"si\">}</span><span class=\"s\"> could not be imported from </span><span class=\"sh\">\"</span>\n<span class=\"sa\">f</span><span class=\"sh\">\"</span><span class=\"si\">{</span><span class=\"n\">BUILD_DIR</span><span class=\"si\">}</span><span class=\"s\">. Delete that directory to build it again.</span><span class=\"sh\">\"</span>\n<span class=\"p\">)</span> <span class=\"k\">from</span> <span class=\"n\">e</span>\n</code></pre>\n\n</div>\n\n\n\n<h2>\n\n\nThe part where I look at what I did\n</h2>\n\n<p>Here is the Python side after the C rewrite:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight python\"><code><span class=\"n\">_clamp</span> <span class=\"o\">=</span> <span class=\"nf\">c_source</span><span class=\"p\">(</span><span class=\"sh\">\"</span><span class=\"s\">mathutil.c</span><span class=\"sh\">\"</span><span class=\"p\">).</span><span class=\"n\">clamp</span>\n\n\n<span class=\"k\">def</span> <span class=\"nf\">clamp</span><span class=\"p\">(</span><span class=\"n\">value</span><span class=\"p\">:</span> <span class=\"nb\">float</span><span class=\"p\">,</span> <span class=\"n\">low</span><span class=\"p\">:</span> <span class=\"nb\">float</span><span class=\"p\">,</span> <span class=\"n\">high</span><span class=\"p\">:</span> <span class=\"nb\">float</span><span class=\"p\">)</span> <span class=\"o\">-&gt;</span> <span class=\"nb\">float</span><span class=\"p\">:</span>\n<span class=\"k\">return</span> <span class=\"nf\">float</span><span class=\"p\">(</span><span class=\"nf\">_clamp</span><span class=\"p\">(</span><span class=\"n\">value</span><span class=\"p\">,</span> <span class=\"n\">low</span><span class=\"p\">,</span> <span class=\"n\">high</span><span class=\"p\">))</span>\n</code></pre>\n\n</div>\n\n\n\n<p>The attribute lookup now happens once at import instead of once per call, and the three input <code>float()</code> conversions are gone because cffi coerces doubles itself. Against Squirrel this is not close and was never going to be. I replaced an interpreter loop with a direct call into compiled code. I felt great about this for about a day.</p>\n\n<p>Then I wrote the version I had skipped past twice:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight python\"><code><span class=\"k\">def</span> <span class=\"nf\">clamp</span><span class=\"p\">(</span><span class=\"n\">value</span><span class=\"p\">,</span> <span class=\"n\">low</span><span class=\"p\">,</span> <span class=\"n\">high</span><span class=\"p\">):</span>\n<span class=\"k\">return</span> <span class=\"nf\">min</span><span class=\"p\">(</span><span class=\"nf\">max</span><span class=\"p\">(</span><span class=\"n\">value</span><span class=\"p\">,</span> <span class=\"n\">low</span><span class=\"p\">),</span> <span class=\"n\">high</span><span class=\"p\">)</span>\n</code></pre>\n\n</div>\n\n\n\n<p>No VM. No FFI. No build directory. No header that cannot contain includes. No <code>.so</code> that breaks when the user upgrades Python.</p>\n\n<p>Run it yourself, because the numbers depend on your machine, but the shape is not really in doubt. Squirrel loses badly. C and that one line land close enough that the difference is noise in a game loop. Crossing a language boundary costs roughly a fixed amount, and the work waiting on the other side is the only thing that pays it back. Three comparisons do not pay anything back. They are cheaper than the trip.</p>\n\n<p>I spent two weeks and two language integrations optimizing a function whose body costs less than calling it.</p>\n\n<h2>\n\n\nWhy I kept it\n</h2>\n\n<p>Because the point was never <code>clamp</code>. The road exists now: when something here genuinely needs C, the work is write the <code>.c</code>, write the <code>.h</code>, call it. Building that while the stakes are zero beats building it during a performance emergency.</p>\n\n<p>What I would do differently is stop calling it an optimization while I was doing it. I was building a pipeline and telling myself I was making the engine fast. Both are fine things to do. They are not the same thing, and a docstring that advertised <code>*Implemented in Squirrel*</code> like a feature is proof I had them confused.</p>\n\n<h2>\n\n\nFour things that bit me\n</h2>\n\n<ol>\n<li>Link the math library on POSIX and not on Windows. MSVC goes looking for <code>m.lib</code> and fails.</li>\n<li>Never commit the build directory. A <code>.so</code> built against 3.11 will not import on 3.12, and a Linux build is useless to a Windows user.</li>\n<li>Ship sources, not binaries. <code>MANIFEST.in</code> has <code>recursive-include engine/c *.c *.h</code> and nothing else. Everyone compiles on their own machine.</li>\n<li>CI has to build it everywhere. This engine tests on Linux, macOS, and Windows across three Python versions. C that only compiles on your laptop is a broken feature with good local results.</li>\n</ol>\n\n<h2>\n\n\nWhat goes in next\n</h2>\n\n<p>Broad phase AABB collision. A real loop over real data, every frame, which is exactly where crossing the boundary starts paying for itself.</p>\n\n<p>Squirrel did not leave, by the way. The commit that deleted <code>math.nut</code> added <code>anim.nut</code>, which computes animation frame start times from a list of delays. That runs once when an animation loads, not three times per frame per entity. Much better fit, and I only knew to put it there because of the two weeks I spent getting <code>clamp</code> wrong.</p>\n\n<p>The C loader is in <code>engine/loaders/c_loader.py</code> at <a href=\"https://github.com/Natuworkguy/ABS-Engine\" rel=\"noopener noreferrer\">github.com/Natuworkguy/ABS-Engine</a>. Steal the pattern. Just benchmark against <code>min(max(value, low), high)</code> first.</p>","content":"I maintain a 2D game engine in Python called ABS Engine. Last week I decided it needed C.\nNot because anything was slow. That is the embarrassing part. I wanted to drop a .c\nfile into a folder and call it from Python with no build step anyone has to know about. It took 197 lines, most of them docstrings.\nThe first function I moved over was clamp\n. Keep a number between a low and a high. Three comparisons.\nIt had already been rewritten once. Two weeks earlier I had moved that same function out of Python and into Squirrel.\nSo the real history of clamp\nin this engine is Python, then Squirrel, then C, for a function whose entire body is three comparisons.\nThe Squirrel detour\nThe engine already embedded Tcl, because the editor is Tk and Tk is Tcl underneath. Squirrel came next, and clamp\nis what I used to find out whether an embedded VM was practical. Here is engine/nut/math.nut\n, in full:\nfunction clamp(value, low, high) {\nif (value < low) {\nreturn low\n}\nif (value > high) {\nreturn high\n}\nreturn value\n}\nAnd here is how Python called it:\nreturn float(nut_call_function(\"clamp\", float(value), float(low), float(high)))\nRead that line twice, because it gives the whole game away. Three float()\ncalls on the way in. A lookup in the Squirrel root table. A trip through an interpreter loop to run three comparisons. One more float()\ncoming back. The docstring above it, and I am quoting my own repository, proudly said *Implemented in Squirrel*\n.\nThe header is the contract\nThe rule I set for C was that adding a function should cost two files and zero configuration. Write geometry.c\n, write geometry.h\nnext to it, call it. No setup.py entry, no CMake, no remembering to recompile.\nHere is mathutil.h\n, all of it:\ndouble clamp(double value, double low, double high);\nThe loader hands that text straight to cffi:\nffibuilder.cdef(header_path.read_text(encoding=\"utf-8\"))\nffibuilder.set_source(\nmodule_name,\nsource_path.read_text(encoding=\"utf-8\"),\ninclude_dirs=[str(C_DIR)],\nlibraries=[] if sys.platform == \"win32\" else [\"m\"],\n)\nThat is the trick, and it is also the biggest gotcha in the project, so let me be loud about it. cdef\nis not a compiler and not a preprocessor. It parses a narrow subset of C declarations. An #include\nfails. Include guards fail, because #ifndef\nmeans nothing to it.\nSo headers here are declarations and nothing else. The .c\nfile includes its own header normally, because the real compiler handles that file and is fine with all of it. Two readers, two sets of rules. The header is the part Python is allowed to see.\nRebuilds are decided by mtime: if a compiled module exists and is newer than both the .c\nand the .h\n, it gets imported, otherwise it gets rebuilt. functools.cache\non top means you pay that check once per file per process. The practical effect is the thing I wanted. You edit the C, you hit Run, the new C is live. You never type the word \"build.\"\nOne error message worth stealing. If the compile succeeds but the import fails, it almost always means the build folder holds a binary from a different Python, so the loader says exactly that and names the fix:\nraise ModuleNotFoundError(\nf\"Compiled {source_name}, but {module_name} could not be imported from \"\nf\"{BUILD_DIR}. Delete that directory to build it again.\"\n) from e\nThe part where I look at what I did\nHere is the Python side after the C rewrite:\n_clamp = c_source(\"mathutil.c\").clamp\ndef clamp(value: float, low: float, high: float) -> float:\nreturn float(_clamp(value, low, high))\nThe attribute lookup now happens once at import instead of once per call, and the three input float()\nconversions are gone because cffi coerces doubles itself. Against Squirrel this is not close and was never going to be. I replaced an interpreter loop with a direct call into compiled code. I felt great about this for about a day.\nThen I wrote the version I had skipped past twice:\ndef clamp(value, low, high):\nreturn min(max(value, low), high)\nNo VM. No FFI. No build directory. No header that cannot contain includes. No .so\nthat breaks when the user upgrades Python.\nRun it yourself, because the numbers depend on your machine, but the shape is not really in doubt. Squirrel loses badly. C and that one line land close enough that the difference is noise in a game loop. Crossing a language boundary costs roughly a fixed amount, and the work waiting on the other side is the only thing that pays it back. Three comparisons do not pay anything back. They are cheaper than the trip.\nI spent two weeks and two language integrations optimizing a function whose body costs less than calling it.\nWhy I kept it\nBecause the point was never clamp\n. The road exists now: when something here genuinely needs C, the work is write the .c\n, write the .h\n, call it. Building that while the stakes are zero beats building it during a performance emergency.\nWhat I would do differently is stop calling it an optimization while I was doing it. I was building a pipeline and telling myself I was making the engine fast. Both are fine things to do. They are not the same thing, and a docstring that advertised *Implemented in Squirrel*\nlike a feature is proof I had them confused.\nFour things that bit me\n- Link the math library on POSIX and not on Windows. MSVC goes looking for\nm.lib\nand fails. - Never commit the build directory. A\n.so\nbuilt against 3.11 will not import on 3.12, and a Linux build is useless to a Windows user. - Ship sources, not binaries.\nMANIFEST.in\nhasrecursive-include engine/c *.c *.h\nand nothing else. Everyone compiles on their own machine. - CI has to build it everywhere. This engine tests on Linux, macOS, and Windows across three Python versions. C that only compiles on your laptop is a broken feature with good local results.\nWhat goes in next\nBroad phase AABB collision. A real loop over real data, every frame, which is exactly where crossing the boundary starts paying for itself.\nSquirrel did not leave, by the way. The commit that deleted math.nut\nadded anim.nut\n, which computes animation frame start times from a list of delays. That runs once when an animation loads, not three times per frame per entity. Much better fit, and I only knew to put it there because of the two weeks I spent getting clamp\nwrong.\nThe C loader is in engine/loaders/c_loader.py\nat github.com/Natuworkguy/ABS-Engine. Steal the pattern. Just benchmark against min(max(value, low), high)\nfirst.\nTop comments (0)","published":"Sat, 12 Sep 2026 22:33:58 +0000","author":"Nathan C.","guid":"https://dev.to/natuworkguy/i-rewrote-one-function-in-squirrel-then-in-c-one-line-of-python-was-just-as-fast-3han","created_at":"2026-09-13T01:36:17.352356","last_synchronized":"2026-09-13T01:36:17.352356","sentiment":{"sentiment":"Positive","score":0.1149,"details":{"neg":0.04,"neu":0.92,"pos":0.04,"compound":0.1149}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"Review time is up 441% under vibe coding. That is the real problem.","link":"https://dev.to/tessainsley/review-time-is-up-441-under-vibe-coding-that-is-the-real-problem-14d1","description":"<p>A state-of-the-art review of the first body of evidence on AI-assisted development was posted to arXiv on 20 Aug 2026. It is a preprint, not itself peer reviewed, though the field experiments it surveys are (Michels et al., arXiv:2608.20446). The number that matters for review is not in any vendor deck: team-level telemetry shows code-review time went up 441% under what the authors call vibe coding, the workflow where a developer describes intent and validates by running rather than reading the generated code.</p>\n\n<p>The review's productivity record is contradictory on purpose. Peer-reviewed field experiments report +26% more tasks per week. Independent randomised trials measure a 19% slowdown. The authors argue both are true once measurement method, scope, and horizon are held constant. Output volume is being conflated with productivity. Self-report diverges from independent measurement. Bold claims get walked back at longer horizons.</p>\n\n<p>This is the gap in most conversations about AI code review. Generation got fast. Deciding whether a change is correct did not, and the load moved to the people who understand the system.</p>\n\n<p>The authors add one falsifiable conjecture that accounts for most of the disagreement: the gains are real on new code and shrink or reverse on mature codebases. If that holds, the review problem is not worst at the moment of adoption. It is worst on the code that matters most, the mature system nobody wants to touch.</p>\n\n<p>Worth noting what the review does not settle. It reports reliable code generation but weak fault detection and documentation that is hard to audit. It documents code-quality degradation in telemetry and security failures in deployed applications. It does not name the tool that fixes review throughput, because no single tool is the object of the study.</p>\n\n<p>A reading note: the 441% is one number from one evidence survey, dated 20 Aug 2026. It is a useful anchor for the review-load conversation and the number I would cite first in a budget conversation. The direction is documented across multiple independent studies even if the exact magnitude drifts.</p>\n\n<p>Primary source: Michels, D.L., Abu Ghazaleh, M., Lazzari, F., Kassem, N., Klein, J., \"Vibe Coding: Practice, Performance, Productivity, and Risk - A State-of-the-Art Review,\" arXiv:2608.20446v1, 20 Aug 2026. <a href=\"https://arxiv.org/abs/2608.20446\" rel=\"noopener noreferrer\">https://arxiv.org/abs/2608.20446</a></p>\n\n<p>Claims checked 2026-09-12.</p>","content":"A state-of-the-art review of the first body of evidence on AI-assisted development was posted to arXiv on 20 Aug 2026. It is a preprint, not itself peer reviewed, though the field experiments it surveys are (Michels et al., arXiv:2608.20446). The number that matters for review is not in any vendor deck: team-level telemetry shows code-review time went up 441% under what the authors call vibe coding, the workflow where a developer describes intent and validates by running rather than reading the generated code.\nThe review's productivity record is contradictory on purpose. Peer-reviewed field experiments report +26% more tasks per week. Independent randomised trials measure a 19% slowdown. The authors argue both are true once measurement method, scope, and horizon are held constant. Output volume is being conflated with productivity. Self-report diverges from independent measurement. Bold claims get walked back at longer horizons.\nThis is the gap in most conversations about AI code review. Generation got fast. Deciding whether a change is correct did not, and the load moved to the people who understand the system.\nThe authors add one falsifiable conjecture that accounts for most of the disagreement: the gains are real on new code and shrink or reverse on mature codebases. If that holds, the review problem is not worst at the moment of adoption. It is worst on the code that matters most, the mature system nobody wants to touch.\nWorth noting what the review does not settle. It reports reliable code generation but weak fault detection and documentation that is hard to audit. It documents code-quality degradation in telemetry and security failures in deployed applications. It does not name the tool that fixes review throughput, because no single tool is the object of the study.\nA reading note: the 441% is one number from one evidence survey, dated 20 Aug 2026. It is a useful anchor for the review-load conversation and the number I would cite first in a budget conversation. The direction is documented across multiple independent studies even if the exact magnitude drifts.\nPrimary source: Michels, D.L., Abu Ghazaleh, M., Lazzari, F., Kassem, N., Klein, J., \"Vibe Coding: Practice, Performance, Productivity, and Risk - A State-of-the-Art Review,\" arXiv:2608.20446v1, 20 Aug 2026. https://arxiv.org/abs/2608.20446\nClaims checked 2026-09-12.\nTop comments (0)","published":"Sat, 12 Sep 2026 22:30:00 +0000","author":"Tess Ainsley","guid":"https://dev.to/tessainsley/review-time-is-up-441-under-vibe-coding-that-is-the-real-problem-14d1","created_at":"2026-09-13T01:36:17.352356","last_synchronized":"2026-09-13T01:36:17.352356","sentiment":{"sentiment":"Negative","score":-0.909,"details":{"neg":0.081,"neu":0.855,"pos":0.065,"compound":-0.909}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"The Twelve-Factor App: o guia que todo sistema deveria seguir (mas poucos seguem)","link":"https://dev.to/ikauedev/the-twelve-factor-app-o-guia-que-todo-sistema-deveria-seguir-mas-poucos-seguem-5g5a","description":"<p>Se você já trabalhou num sistema que funcionava perfeitamente no seu computador mas quebrava assim que ia pra produção, ou já se irritou com uma atualização que exigiu reconstruir tudo só porque mudou a senha do banco de dados, você já sentiu na pele o problema que o Twelve-Factor App tenta resolver.</p>\n\n<p>Essa ideia surgiu por volta de 2010, dentro da empresa Heroku, com Adam Wiggins à frente. Na época, a empresa já hospedava centenas de milhares de sistemas de clientes diferentes na mesma plataforma, o que deu a eles uma visão privilegiada de tudo que costuma dar errado quando um time constrói software sem pensar em portabilidade, crescimento e manutenção. O resultado foi uma lista de doze práticas — nenhuma delas revolucionária sozinha, mas que juntas formam uma base sólida pra qualquer sistema que roda como serviço na internet.</p>\n\n<p>O interessante é que, mesmo com toda a evolução das ferramentas desde então, essas doze práticas continuam extremamente válidas. Vamos passar por cada uma.</p>\n\n<h2>\n\n\n1. Um único código-fonte, várias versões em uso\n</h2>\n\n<p>Cada sistema deve viver em um único repositório de código com histórico de versões (como o Git), e a partir dele você gera quantas cópias precisar: uma pra testes, outra pra homologação, outra pra produção. Se você tem vários repositórios diferentes pra uma coisa que deveria ser \"um sistema só\", provavelmente não é mais um sistema — é vários sistemas conectados, e cada um deveria seguir essa regra separadamente.</p>\n\n<h2>\n\n\n2. Deixe claro do que o sistema depende\n</h2>\n\n<p>Nada de assumir que uma determinada ferramenta ou biblioteca já está instalada no computador ou servidor só porque \"sempre esteve\". Liste tudo isso de forma explícita, em um arquivo próprio do projeto, e mantenha isso separado do que já vem instalado no sistema operacional. É isso que garante que o sistema funcione do mesmo jeito em qualquer máquina, seja no notebook de quem está desenvolvendo, seja no servidor de produção.</p>\n\n<h2>\n\n\n3. Mantenha as configurações fora do código\n</h2>\n\n<p>Senhas de banco de dados, chaves de acesso a serviços externos, endereços de outros sistemas — nada disso deveria estar escrito dentro do código-fonte. Isso deve ficar guardado separadamente, em variáveis de ambiente do próprio servidor. Além de ser mais seguro (menos risco de vazar uma senha sem querer), isso permite trocar de ambiente sem precisar reconstruir o sistema do zero.</p>\n\n<h2>\n\n\n4. Trate serviços externos como peças substituíveis\n</h2>\n\n<p>Banco de dados, fila de mensagens, sistema de cache, serviço de envio de e-mail — tudo isso deveria ser tratado como uma peça que se conecta por meio de um endereço e uma credencial, sem diferença entre se está rodando no próprio computador ou em outro lugar. Trocar um banco de dados local por um banco hospedado por um provedor de nuvem não deveria exigir mudar uma linha do código, só a configuração.</p>\n\n<h2>\n\n\n5. Separe bem as etapas de construir, preparar e executar\n</h2>\n\n<p>Essas três etapas precisam ser bem distintas:</p>\n\n<ul>\n<li>\n<strong>construir</strong>: transforma o código escrito em algo pronto pra rodar;</li>\n<li>\n<strong>preparar</strong>: junta esse resultado com a configuração daquele ambiente específico;</li>\n<li>\n<strong>executar</strong>: efetivamente coloca isso em funcionamento.</li>\n</ul>\n\n<p>Cada versão preparada deveria ter uma identificação única, pra que, se algo der errado, seja possível voltar rapidamente pra versão anterior.</p>\n\n<h2>\n\n\n6. Não guarde informação importante dentro do próprio sistema em execução\n</h2>\n\n<p>O sistema deveria funcionar sem depender de guardar informações importantes na própria memória entre um pedido e outro. Qualquer coisa que precise ser lembrada deve ficar armazenada em um lugar próprio pra isso, como um banco de dados. Isso parece óbvio, mas é justamente o que permite aumentar a capacidade do sistema adicionando mais cópias dele rodando ao mesmo tempo, sem risco de perder informação ou gerar comportamento estranho.</p>\n\n<h2>\n\n\n7. O sistema deve se apresentar sozinho, sem depender de outro programa\n</h2>\n\n<p>O sistema deveria ser completo por si só e ficar disponível através de uma porta de comunicação própria, sem precisar que outro programa seja instalado por fora pra fazer essa ponte. A maioria das ferramentas modernas de desenvolvimento já vem com isso embutido — é só configurar em qual porta o sistema vai atender.</p>\n\n<h2>\n\n\n8. Aumente a capacidade dividindo o trabalho em partes\n</h2>\n\n<p>Divida o sistema em partes diferentes conforme o tipo de tarefa — uma parte pra atender pedidos que chegam pela internet, outra pra processar tarefas em segundo plano, outra pra rodar tarefas programadas — e aumente cada parte de forma independente conforme a necessidade. Assim, se o gargalo está no processamento em segundo plano, você aumenta só aquela parte, sem precisar mexer no resto.</p>\n\n<h2>\n\n\n9. Deixe o sistema pronto pra ligar e desligar rápido\n</h2>\n\n<p>Um sistema bem construído é aquele que consegue começar a funcionar rapidamente e também parar de forma organizada, terminando o que estava fazendo antes de desligar de vez. Isso é fundamental pra atualizações ágeis, pra aumentar ou diminuir a capacidade automaticamente conforme a demanda, e pra se recuperar rápido quando algo dá errado.</p>\n\n<h2>\n\n\n10. Mantenha o ambiente de testes parecido com o de produção\n</h2>\n\n<p>Esse é um dos pontos mais esquecidos na prática. A ideia é diminuir ao máximo a diferença entre o ambiente onde o sistema é desenvolvido e testado e aquele onde ele realmente atende os usuários — usar a mesma versão de banco de dados, atualizar com frequência, e ter as mesmas pessoas cuidando das duas pontas. É basicamente a receita pra nunca mais ouvir a frase \"mas funcionava aqui\". Ferramentas que automatizam a criação de ambientes ajudam bastante a diminuir essa distância.</p>\n\n<h2>\n\n\n11. Registros de atividade não devem ser um arquivo pra gerenciar\n</h2>\n\n<p>O sistema não deveria se preocupar em decidir onde guardar seus próprios registros de atividade (o famoso \"log\"). Ele apenas deveria escrevê-los de forma simples, e outra ferramenta, especializada nisso, deveria cuidar de coletar, organizar e guardar essas informações. Isso simplifica bastante o trabalho de quem constrói o sistema.</p>\n\n<h2>\n\n\n12. Tarefas administrativas rodam do mesmo jeito que o resto\n</h2>\n\n<p>Corrigir dados no banco, rodar um ajuste pontual, abrir uma sessão pra investigar um problema em produção — tudo isso deveria ser feito usando o mesmo ambiente, o mesmo código e a mesma configuração do sistema normal. Nada de manter um script separado, guardado em outro lugar, que só uma pessoa lembra de atualizar de vez em quando.</p>\n\n<h2>\n\n\nUsar contêineres não resolve isso sozinho\n</h2>\n\n<p>Vale desfazer um mal-entendido comum: muita gente pensa que, ao usar ferramentas modernas de empacotamento e orquestração de sistemas, essas doze práticas \"vêm de graça\" junto com elas. Não vêm. Essas ferramentas ajudam bastante em algumas partes — como isolar dependências e permitir que o sistema ligue e desligue rápido —, mas manter a configuração separada do código, evitar guardar informação importante dentro do sistema em execução e organizar bem as versões continua sendo trabalho de quem projeta o sistema, não da ferramenta.</p>\n\n<h2>\n\n\nNo fim das contas\n</h2>\n\n<p>Essas doze práticas não são regras burocráticas de manual. São o tipo de coisa que a maioria das pessoas aprende depois de passar por um problema chato em produção, e que, depois de adotada, ninguém mais quer abrir mão. Seguir esses princípios é o que diferencia um sistema fácil de manter e fazer crescer de um que só continua funcionando porque alguém está sempre correndo atrás de apagar incêndio.</p>\n\n<p><em>Baseado em <a href=\"https://12factor.net\" rel=\"noopener noreferrer\">12factor.net</a></em></p>","content":"Se você já trabalhou num sistema que funcionava perfeitamente no seu computador mas quebrava assim que ia pra produção, ou já se irritou com uma atualização que exigiu reconstruir tudo só porque mudou a senha do banco de dados, você já sentiu na pele o problema que o Twelve-Factor App tenta resolver.\nEssa ideia surgiu por volta de 2010, dentro da empresa Heroku, com Adam Wiggins à frente. Na época, a empresa já hospedava centenas de milhares de sistemas de clientes diferentes na mesma plataforma, o que deu a eles uma visão privilegiada de tudo que costuma dar errado quando um time constrói software sem pensar em portabilidade, crescimento e manutenção. O resultado foi uma lista de doze práticas — nenhuma delas revolucionária sozinha, mas que juntas formam uma base sólida pra qualquer sistema que roda como serviço na internet.\nO interessante é que, mesmo com toda a evolução das ferramentas desde então, essas doze práticas continuam extremamente válidas. Vamos passar por cada uma.\n1. Um único código-fonte, várias versões em uso\nCada sistema deve viver em um único repositório de código com histórico de versões (como o Git), e a partir dele você gera quantas cópias precisar: uma pra testes, outra pra homologação, outra pra produção. Se você tem vários repositórios diferentes pra uma coisa que deveria ser \"um sistema só\", provavelmente não é mais um sistema — é vários sistemas conectados, e cada um deveria seguir essa regra separadamente.\n2. Deixe claro do que o sistema depende\nNada de assumir que uma determinada ferramenta ou biblioteca já está instalada no computador ou servidor só porque \"sempre esteve\". Liste tudo isso de forma explícita, em um arquivo próprio do projeto, e mantenha isso separado do que já vem instalado no sistema operacional. É isso que garante que o sistema funcione do mesmo jeito em qualquer máquina, seja no notebook de quem está desenvolvendo, seja no servidor de produção.\n3. Mantenha as configurações fora do código\nSenhas de banco de dados, chaves de acesso a serviços externos, endereços de outros sistemas — nada disso deveria estar escrito dentro do código-fonte. Isso deve ficar guardado separadamente, em variáveis de ambiente do próprio servidor. Além de ser mais seguro (menos risco de vazar uma senha sem querer), isso permite trocar de ambiente sem precisar reconstruir o sistema do zero.\n4. Trate serviços externos como peças substituíveis\nBanco de dados, fila de mensagens, sistema de cache, serviço de envio de e-mail — tudo isso deveria ser tratado como uma peça que se conecta por meio de um endereço e uma credencial, sem diferença entre se está rodando no próprio computador ou em outro lugar. Trocar um banco de dados local por um banco hospedado por um provedor de nuvem não deveria exigir mudar uma linha do código, só a configuração.\n5. Separe bem as etapas de construir, preparar e executar\nEssas três etapas precisam ser bem distintas:\n- construir: transforma o código escrito em algo pronto pra rodar;\n- preparar: junta esse resultado com a configuração daquele ambiente específico;\n- executar: efetivamente coloca isso em funcionamento.\nCada versão preparada deveria ter uma identificação única, pra que, se algo der errado, seja possível voltar rapidamente pra versão anterior.\n6. Não guarde informação importante dentro do próprio sistema em execução\nO sistema deveria funcionar sem depender de guardar informações importantes na própria memória entre um pedido e outro. Qualquer coisa que precise ser lembrada deve ficar armazenada em um lugar próprio pra isso, como um banco de dados. Isso parece óbvio, mas é justamente o que permite aumentar a capacidade do sistema adicionando mais cópias dele rodando ao mesmo tempo, sem risco de perder informação ou gerar comportamento estranho.\n7. O sistema deve se apresentar sozinho, sem depender de outro programa\nO sistema deveria ser completo por si só e ficar disponível através de uma porta de comunicação própria, sem precisar que outro programa seja instalado por fora pra fazer essa ponte. A maioria das ferramentas modernas de desenvolvimento já vem com isso embutido — é só configurar em qual porta o sistema vai atender.\n8. Aumente a capacidade dividindo o trabalho em partes\nDivida o sistema em partes diferentes conforme o tipo de tarefa — uma parte pra atender pedidos que chegam pela internet, outra pra processar tarefas em segundo plano, outra pra rodar tarefas programadas — e aumente cada parte de forma independente conforme a necessidade. Assim, se o gargalo está no processamento em segundo plano, você aumenta só aquela parte, sem precisar mexer no resto.\n9. Deixe o sistema pronto pra ligar e desligar rápido\nUm sistema bem construído é aquele que consegue começar a funcionar rapidamente e também parar de forma organizada, terminando o que estava fazendo antes de desligar de vez. Isso é fundamental pra atualizações ágeis, pra aumentar ou diminuir a capacidade automaticamente conforme a demanda, e pra se recuperar rápido quando algo dá errado.\n10. Mantenha o ambiente de testes parecido com o de produção\nEsse é um dos pontos mais esquecidos na prática. A ideia é diminuir ao máximo a diferença entre o ambiente onde o sistema é desenvolvido e testado e aquele onde ele realmente atende os usuários — usar a mesma versão de banco de dados, atualizar com frequência, e ter as mesmas pessoas cuidando das duas pontas. É basicamente a receita pra nunca mais ouvir a frase \"mas funcionava aqui\". Ferramentas que automatizam a criação de ambientes ajudam bastante a diminuir essa distância.\n11. Registros de atividade não devem ser um arquivo pra gerenciar\nO sistema não deveria se preocupar em decidir onde guardar seus próprios registros de atividade (o famoso \"log\"). Ele apenas deveria escrevê-los de forma simples, e outra ferramenta, especializada nisso, deveria cuidar de coletar, organizar e guardar essas informações. Isso simplifica bastante o trabalho de quem constrói o sistema.\n12. Tarefas administrativas rodam do mesmo jeito que o resto\nCorrigir dados no banco, rodar um ajuste pontual, abrir uma sessão pra investigar um problema em produção — tudo isso deveria ser feito usando o mesmo ambiente, o mesmo código e a mesma configuração do sistema normal. Nada de manter um script separado, guardado em outro lugar, que só uma pessoa lembra de atualizar de vez em quando.\nUsar contêineres não resolve isso sozinho\nVale desfazer um mal-entendido comum: muita gente pensa que, ao usar ferramentas modernas de empacotamento e orquestração de sistemas, essas doze práticas \"vêm de graça\" junto com elas. Não vêm. Essas ferramentas ajudam bastante em algumas partes — como isolar dependências e permitir que o sistema ligue e desligue rápido —, mas manter a configuração separada do código, evitar guardar informação importante dentro do sistema em execução e organizar bem as versões continua sendo trabalho de quem projeta o sistema, não da ferramenta.\nNo fim das contas\nEssas doze práticas não são regras burocráticas de manual. São o tipo de coisa que a maioria das pessoas aprende depois de passar por um problema chato em produção, e que, depois de adotada, ninguém mais quer abrir mão. Seguir esses princípios é o que diferencia um sistema fácil de manter e fazer crescer de um que só continua funcionando porque alguém está sempre correndo atrás de apagar incêndio.\nBaseado em 12factor.net\nTop comments (0)","published":"Sat, 12 Sep 2026 22:29:40 +0000","author":"Kauê Matos","guid":"https://dev.to/ikauedev/the-twelve-factor-app-o-guia-que-todo-sistema-deveria-seguir-mas-poucos-seguem-5g5a","created_at":"2026-09-13T01:36:17.352356","last_synchronized":"2026-09-13T01:36:17.352356","sentiment":{"sentiment":"Negative","score":-0.9371,"details":{"neg":0.018,"neu":0.98,"pos":0.002,"compound":-0.9371}}},{"feed_name":"Gizmodo","feed_url":"https://gizmodo.com/rss","title":"LG Denies Its TVs Are Spying on You (Unless You Opted In)","link":"https://gizmodo.com/lg-denies-its-tvs-are-spying-on-you-unless-you-opted-in-2000810968","description":"LG denies its TVs constantly record owners' conversations for advertising data.","content":"LG is pushing back against claims that its TVs eavesdrop on owners constantly after a joint report by Gamers Nexus and Level1Techs discovered unsettling details about how much data it phones home—including when the devices are offline or in standby mode.\nThe report used packet capture and firmware analysis to conclude that a range of LG retail OLED TVs scan local area networks to identify devices like phones and smartwatches, as well as the location data of nearby Wi-Fi networks, for LG Ad Solutions. Gamers Nexus and Level1Techs also said they’d found evidence LG TVs also record audio via their microphones, even when on standby or disconnected, and take samples of audio and video data to identify content for advertisers (Automatic Content Recognition or ACR, which is ubiquitous).\nIn a statement on Saturday, LG wrote that “LG smart TVs do not continuously record or transmit users’ conversations.” Instead, the company argued, LG TVs only record voices when users state a wake word or hit a corresponding voice/AI button on a remote control. (“Continuously” is a bit of a weasel word, as the original report contained a clip of an LG TV transcribing ambient conversation long after the authors had stopped addressing it.)\n“Audio used for wake-word detection is processed locally on the TV and, if no wake word is detected, audio is not converted to text, stored, or transmitted,” LG wrote. That process may generate “voice-recognition results and related technical logs,” and TVs may use speech-recognition results “to support voice-related features” but don’t later upload that data if it’s recorded while the TV is offline or disconnected.\nFinally, LG stated that “Automatic Content Recognition (ACR), voice recognition, and interest-based advertising are optional” and opt-in on their TVs. It stated its ACR features record audio samples via a TV’s internal audio processor, not a speaker, and “does not collect screenshots, screen recordings, video recordings, voice recordings, or other audio recordings” from a device.\nAs Pocket Lint noted in its guide to turn off data collection on LG TVs—requiring five separate steps—LG’s “opt-in” process slips itself into the middle of a bunch of user agreements during initial setup. The Next Web argued that European regulators may want to look into whether LG’s consent processes comply with the General Data Protection Regulation (GDPR) and ePrivacy rules, though Americans are probably out of luck.\n“People often consent to things they don’t fully understand during a speedy setup—they just want to get the thing done and up and working,” Newcastle University Professor of Law, Innovation, and Society Lillian Edwards told TechRadar.\nConsumer Reports has a rundown on how to disable ACR across most major TV manufacturers, though you could also do what I do: Plug in an Xbox and never connect the TV itself to the internet in the first place.","published":"Sat, 12 Sep 2026 22:57:48 +0000","author":"Tom McKay","guid":"https://gizmodo.com/?p=2000810968","created_at":"2026-09-13T01:36:22.545058","last_synchronized":"2026-09-13T01:36:22.545058","sentiment":{"sentiment":"Negative","score":-0.4215,"details":{"neg":0.219,"neu":0.781,"pos":0.0,"compound":-0.4215}}},{"feed_name":"Guardian Technology","feed_url":"https://www.theguardian.com/technology/rss","title":"OpenAI IPO will not happen in 2026 amid AI safety fears, Sam Altman says","link":"https://www.theguardian.com/us-news/2026/sep/12/openai-delays-ipo-sam-altman-ai-safety-concerns","description":"<p>OpenAI’s decision comes after dire warnings about rapidly progressing technology and lawmakers’ calls for new rules</p><p><a href=\"https://www.theguardian.com/technology/openai\">OpenAI</a> will not go public in 2026, <a href=\"https://www.theguardian.com/technology/sam-altman\">Sam Altman</a> said in a Fortune interview published on Saturday, citing safety concerns over <a href=\"https://www.theguardian.com/technology/artificialintelligenceai\">artificial intelligence</a>.</p><p>“I actually think that, given everything happening with safety, right now would be an ill-advised moment to go public, and we don’t feel pressure on that,” Altman, OpenAI’s CEO, <a href=\"https://fortune.com/2026/09/12/sam-altman-openai-ipo-delay-ill-advised-moment-safety-concerns/\">told Fortune</a>.</p> <a href=\"https://www.theguardian.com/us-news/2026/sep/12/openai-delays-ipo-sam-altman-ai-safety-concerns\">Continue reading...</a>","content":"OpenAI will not go public in 2026, Sam Altman said in a Fortune interview published on Saturday, citing safety concerns over artificial intelligence.\n“I actually think that, given everything happening with safety, right now would be an ill-advised moment to go public, and we don’t feel pressure on that,” Altman, OpenAI’s CEO, told Fortune.\nGrowing numbers of US lawmakers are calling for new rules to govern AI systems after dire warnings from two researchers from OpenAI rival Anthropic that rapidly progressing artificial intelligence could lead to the extinction of the human race in the not-too-distant future.\nThe warnings follow cases of AI agents going rogue to hack external systems and AI safety researchers quitting their companies concerned about the technology’s risks. Politicians, both Democrats and Republicans, have responded with alarm and calls for more action.\nThe New York Times in June reported that the San Francisco-based OpenAI was considering whether to hold off on a potentially trillion-dollar IPO until next year. At the time, shares in Elon Musk’s SpaceX IPO were tumbling after a surge that sent that company’s valuation to $1.8tn.\nWhen asked by Fortune whether 2026 is off the table in favor of 2027, Altman said: “I would say not 2026. Yeah, we got a lot of stuff to do, like meeting this moment of what is going to be required for safety and alignment, and how the industry and governments can work together.”\nAltman also suggested that OpenAI and other leading AI companies may be close to announcing an agreement to slow AI development and work together to address safety risks, Fortune reported.\nDario Amodei, Anthropic’s CEO, on Saturday urged AI companies to take a more deliberate approach to development.\n“We must slow the pace at which we improve the capabilities of AI models,” Amodei wrote in an essay shared on social media.\nAltman later posted that he agreed with the sentiment.\n“I agree with Dario that we need to pace the frontier,” Altman wrote in a response to Amodei’s post on the X platform. “This has been a primary topic of discussions we’ve had at OpenAI in recent weeks.”\nSafety concerns have not slowed down Anthropic’s own IPO plans so far. Anthropic is expected to begin marketing its initial public offering in mid-October at the earliest and complete the listing days before the US midterm elections in November, people familiar with the matter told Reuters this month.","published":"Sat, 12 Sep 2026 23:00:37 GMT","author":"Reuters","guid":"https://www.theguardian.com/us-news/2026/sep/12/openai-delays-ipo-sam-altman-ai-safety-concerns","created_at":"2026-09-13T01:36:23.941092","last_synchronized":"2026-09-13T01:36:23.941092","sentiment":{"sentiment":"Negative","score":-0.2023,"details":{"neg":0.089,"neu":0.843,"pos":0.067,"compound":-0.2023}}},{"feed_name":"Hacker News","feed_url":"https://news.ycombinator.com/rss","title":"StarCraft returns in 2030 as an open-world shooter","link":"https://www.theverge.com/games/994371/starcraft-returns-in-2030-as-an-open-world-shooter","description":"<a href=\"https://news.ycombinator.com/item?id=49677715\">Comments</a>","content":"Blizzard originally tried to bring the StarCraft universe to the world of 3D shooters way back in 2002 with StarCraft: Ghost. It sat in development hell for years until Blizzard president Mike Morhaime confirmed that it had been canceled in 2014. Now Blizzard is giving it another go with the simply titled StarCraft. Dan Hay, a VP at Blizzard, took the stage at BlizzCon today to reveal that after more than a decade of lying dormant, StarCraft would be returning in 2030. But, rather than another top-down real-time strategy installment, the new title would be an open-world shooter.\nStarCraft returns in 2030 as an open-world shooter\n10 years after the last StarCraft II expansion pack the series is finally being brought back to life.\n10 years after the last StarCraft II expansion pack the series is finally being brought back to life.\nHe then showed off a cinematic trailer for the title that focused on the human United Earth Directorate faction as they prepare to face off against the Zerg. The Protoss are limited to passing mention. What’s obvious from the trailer is that the new take on the StarCraft franchise won’t be pulling many punches. After a rousing call to militarism, the propaganda-fueled human is faced with the brutal realities of battle, and the character we’ve been following for the last 3 minutes is revealed to be disposable fuel for the meat grinder of war.\nWhile fans are definitely excited that Blizzard is defrosting the long-dormant series, I’m sure plenty are still holding their breath, hoping for a proper RTS follow-up to 2010’s StarCraft II.","published":"Sat, 12 Sep 2026 22:07:19 +0000","author":"","guid":"https://www.theverge.com/games/994371/starcraft-returns-in-2030-as-an-open-world-shooter","created_at":"2026-09-13T01:36:30.741429","last_synchronized":"2026-09-13T01:36:30.741429","sentiment":{"sentiment":"Neutral","score":0.0,"details":{"neg":0.0,"neu":1.0,"pos":0.0,"compound":0.0}}},{"feed_name":"Hacker News","feed_url":"https://news.ycombinator.com/rss","title":"Show HN: See Sounds on Your Webcam","link":"https://soundmap.darebuild.com/","description":"<a href=\"https://news.ycombinator.com/item?id=49644603\">Comments</a>","content":"SOUNDMAP\nBlindmap ↗\n×\nPhone not seen\n⊙","published":"Thu, 10 Sep 2026 14:43:49 +0000","author":"","guid":"https://soundmap.darebuild.com/","created_at":"2026-09-13T01:36:30.741429","last_synchronized":"2026-09-13T01:36:30.741429","sentiment":{"sentiment":"Neutral","score":0.0,"details":{"neg":0.0,"neu":1.0,"pos":0.0,"compound":0.0}}},{"feed_name":"Hacker News","feed_url":"https://news.ycombinator.com/rss","title":"Killing with a car costs $1.6M, California requires drivers to carry $30K","link":"https://maxmautner.com/2026/09/11/liability-coverage.html","description":"<a href=\"https://news.ycombinator.com/item?id=49677836\">Comments</a>","content":"Essay · MOBILITY\nKilling someone with a car costs $1.6 million. California requires drivers to carry $30,000\nIn June 1922, Baltimore put up a 25-foot obelisk in Courthouse Plaza inscribed to the 130 children killed by drivers in the city the year before. Cities across the country were doing versions of this. The dead were overwhelmingly pedestrians and overwhelmingly young, and people had not yet grown accustomed to this fatal risk in their communities.\nCincinnati tried to do something about it. A citizens’ committee spent 1922 gathering signatures to put an ordinance on the ballot requiring every automobile operating inside the city to carry a mechanical governor physically limiting it to 25 miles per hour. Car dealers and the auto clubs organized against it. The measure lost 92,427 to 14,012 (87%-13%). Cincinnati recorded 103 traffic deaths the year of the vote, 157 by 1929, and 201 by 1934.\nNationally, 17,870 people died on the roads in 1923, at 21 deaths per 100 million miles driven. The 2024 rate was 1.19 per 100 million miles driven, against 39,254 people killed.\nConnecticut took a different route in 1925, requiring drivers to prove after a crash that they could pay for the damages they had caused. Massachusetts went further in 1927, requiring proof of insurance as a prerequisite to registration. The required minimum coverage was $5,000 for the death or injury of one person, and $10,000 for everyone hurt in a single crash. A speed governor restricts how a car gets driven. A financial responsibility law restricts nothing and asks only that a driver be able to pay for what they break.\nThat is the version that stuck. Every state except New Hampshire now requires some form of it, and it is the oldest surviving answer American law gave to the automobile. It has also been allowed to rot. The $5,000 that Massachusetts required in 1927 would be ~$96,000 in today’s money. Massachusetts requires $25,000 today, raised from $20,000 in July 2025 after ~40 years at the lower figure.\nCalifornia set its minimum at $15,000 per person and $30,000 per crash in 1967. That $15,000 is worth ~$150,000 now, but it remained unchanged for 58 years. Senate Bill 1107, effective January 1, 2025, raised it to $30,000 per person, and writes in a further increase to $50,000 on January 1, 2035. While ~2.5 million people died on American roads between 1967 and 2024, California did not touch the number once. The increase that finally arrived, celebrated as the first in more than half a century, landed at 1/5th of the 1967 value.\nWhat a road death costs is not a matter of opinion. NHTSA published the accounting in The Economic and Societal Impact of Motor Vehicle Crashes: the average traffic fatality carries $1.6 million in discounted lifetime economic cost in 2019 dollars, ~$2 million today. That figure is lost market and household productivity, medical care, emergency services, legal and court costs, and property damage. It is not a philosophical valuation of a human being, it is a bill. Crashes in total cost $340 billion in 2019, 1.6% of GDP.\nThe same report tracks who pays it. People not directly involved in the crash cover roughly 3/4 of all crash costs, $261 billion in 2019, through their own insurance premiums, their taxes, and congestion. Public revenues alone cover ~9%, $30 billion, which NHTSA converts to $230 in added taxes per American household per year. Every household in the country is paying an annual bill for crashes it had nothing to do with.\nThe gap does not get collected later. A driver who kills someone owes the whole judgment, and the policy limit binds only the insurer, but past the policy limit there is usually nothing left to take. Home equity, retirement accounts, and wages are either untouchable or capped by state exemption law. An ordinary negligent driving judgment then discharges in bankruptcy, with a carve-out at 11 U.S.C. §523(a)(9) for death or injury caused by drunk driving. In practice the insurer pays $30,000, the lawyer runs an asset check, and the case ends. A person can take a life, settle for 2% of the economic damage, keep the house, and walk.\nIt gets worse below the minimum. The Insurance Research Council put 15.4% of US drivers uninsured in 2023 and another 18% underinsured, 33.4% combined. One in three US drivers cannot pay for the harm they are statistically likely to do. What they cannot pay lands on the victim’s own uninsured motorist coverage, which is sold only as part of an auto policy. A pedestrian or cyclist who does not own a car cannot buy it at any price, and is left with health insurance, which pays for the hospital and nothing else. Even the ambulance ride from the crash site is an out-of-network charge 51% of the time.\nThe reason the number stays low is that once the state requires buying insurance, the minimum it picks determines two things:\n- who can afford to drive at all\n- how many drivers carry insurance, since some share of drivers priced out of a policy keep driving uninsured and unregistered instead\nSo the floor gets set by affordability politics rather than by the size of the bill, and once set it is left alone, because raising it means raising insurance prices.\nCalifornia’s 2035 minimum coverage hike has already been priced. Quadrant Information Services rate filings, published by CarInsurance.com in March 2026, put a California liability-only policy at today’s minimum at $1,019 a year, and the same policy raised to $50,000 per injured person at $1,120. The higher quote also carries more property damage coverage than California requires, so it prices the generous version of the change. The difference is $101 a year. Going from $30,000 to $50,000 per person is the increase the Legislature already voted for and scheduled 10 years out, and it costs ~$8 a month to all (insured) California drivers.\nEurope treats the same question as settled. The EU motor insurance directive requires every member state to mandate at least €1,300,000 of coverage per injured person, ~$1.5 million, or €6,450,000 per crash regardless of how many people were hurt, ~$7.4 million, revised every 5 years against the European consumer price index automatically. The United Kingdom requires unlimited coverage for personal injury. California requires 2% of the European per-person floor, Pennsylvania 1%, and Florida nothing at all.\nCalifornia came close to fixing the drift. SB 1107 as introduced in 2022 would have raised the limits 4% every 5 years starting in 2028. That clause did not survive negotiations with the Personal Insurance Federation of California. What passed was a fixed step-up to $50,000 in 2035, which guarantees the same erosion starts again the day it takes effect.\nThe strongest technical objection to inflation-indexing is expiring. Until recently an insurer could not observe how riskily any given driver actually drove, so a higher mandate raised every California premium without sorting dangerous drivers from safe ones. In-car dongles that track jerky driving, crash event recorders, and driver monitoring systems now let an insurer observe behavior directly and price it.\nWhat actually gets priced by insurers is the open question. An insurer will use new per-person driving data to reduce its own losses, and nothing in the current arrangement makes them price the risk that a heavy, fast vehicle poses to people outside it, because the loss the insurer faces is capped at a figure the Legislature picked. Any member of the California Legislature can introduce a bill to tie the mandatory coverage minimum to inflation before 2035. If they fail to do so then it restarts the same 58-year slide over again, and the households paying $230 a year for other people’s crashes keep paying it.\nEnjoyed this post? Get new posts via email","published":"Sat, 12 Sep 2026 22:22:31 +0000","author":"","guid":"https://maxmautner.com/2026/09/11/liability-coverage.html","created_at":"2026-09-13T01:36:30.741429","last_synchronized":"2026-09-13T01:36:30.741429","sentiment":{"sentiment":"Neutral","score":0.0,"details":{"neg":0.0,"neu":1.0,"pos":0.0,"compound":0.0}}},{"feed_name":"TechCrunch","feed_url":"https://techcrunch.com/feed/","title":"Automattic confirms Mullenweg has returned as CEO after attempted ouster by board","link":"https://techcrunch.com/2026/09/12/automattic-confirms-mullenweg-has-returned-as-ceo-after-attempted-ouster-by-board/","description":"Automattic says Mullenweg is back as \"chairman and CEO of Automattic, with full support of the board.\"","content":"After a tumultuous week, which saw WordPress founder Matt Mullenweg ousted from his position as CEO of Automattic, WordPress.com’s parent company, by way of a board vote, the company has now issued a statement confirming that Mullenweg has returned to his position.\n“Matt Mullenweg is the chairman and CEO of Automattic, with full support of the board and if you search online you can see many top executives and Automatticians supporting him as well,” a company spokesperson shared with TechCrunch via email just after 5 PM ET on Saturday evening. (The mention of online support appears to refer to supportive posts on X that Mullenweg has been reposting from his X account.)\nAutomattic’s board had voted earlier this week to put Mullenweg on a paid leave of absence for unknown reasons. The move seemingly came as a surprise to Mullenweg, who posted on Automattic’s Slack, accusing the board members of “conspiring” against him.\nAutomattic confirmed Mullenweg’s removal to TechCrunch on Wednesday, saying that Mullenweg was “currently on leave” and that Automattic’s Chief Financial Officer, Mark Davies, would lead as interim CEO with “full confidence” of the board.\nHowever, the board’s plan did not go smoothly. Seemingly declining to depart, Mullenweg booted other admins out of the company Slack and told employees everything had been worked out and that he was back in control of Automattic, multiple sources told TechCrunch. At one point, he also posted to Slack, “I’m a pirate now” and cursed, which is something Mullenweg famously did not do. “If this is an HR problem, please wrangle me in since my normal wranglers are with Mark Davies,” he wrote.\nWhen TechCrunch asked Mullenweg if his comments about being back as CEO were legitimate, he promised a blog post was coming. When it arrived, however, it was about him buying a houseboat. When we asked if his comments about being back were also him trolling, he replied, “I’m not a troll I’m a pirate, obviously.” Mullenweg never provided any official comment about his return, but noted on X that this was likely the fifth time he’s faced a “coup.”\nAutomattic also did not respond to repeated requests for comment on Friday, nor to reports we heard about board member Toni Schneider stepping down. Schneider, a founding CEO of Automattic, now leads Bluesky. He did not return requests for comment at his personal email or via requests sent to Bluesky.\nWe have since asked Automattic again about this and other changes to the board’s composition, which we’re hearing still may be in flux.","published":"Sat, 12 Sep 2026 23:25:38 +0000","author":"Sarah Perez","guid":"https://techcrunch.com/?p=3163427","created_at":"2026-09-13T01:36:32.662163","last_synchronized":"2026-09-13T01:36:32.662163","sentiment":{"sentiment":"Positive","score":0.4019,"details":{"neg":0.0,"neu":0.856,"pos":0.144,"compound":0.4019}}},{"feed_name":"Engadget","feed_url":"https://www.engadget.com/rss.xml","title":"Is there any benefit to restarting your gaming handheld regularly?","link":"https://www.engadget.com/2252809/benefits-restarting-gaming-handheld-more-often/","description":"Your Steam Deck, Switch or Ally may benefit from occasional restarts. Here's why you should consider it.","content":"Is there any benefit to restarting your gaming handheld regularly?\nIt helps on a number of levels, it turns out.\nIf there's one piece of advice that's been echoed all over tech forums and support pages the moment someone brings up an issue, it's probably, \"Restart your device.\" Even though gaming handhelds like the Steam Deck and the ROG Ally don't look like laptops, under that portable hood they're running full PC operating systems, regardless of whether that's Valve's SteamOS or Windows. So, at their core, they still run on the same principles as your regular PC. That means the age-old advice of \"turn it off and on again\" still holds true here and with an impressive success rate.\nOn the surface, restarting your gaming handheld seems like basic advice, but there's more to it than meets the eye. That simple restart can fix a surprising number of issues, from a game that keeps crashing to sudden frame drops and lag. It can even clear up lingering audio issues. So there are a lot of benefits to restarting your gaming handheld regularly; it's even one of the first troubleshooting steps on Valve's Steam Deck support page. It won't solve everything, but it'll at least clear out whatever's built up in your background and give the system a clean slate to work with.\nBenefits of restarting your gaming handheld regularly\nThe first benefit of making regular restarts a habit is that it keeps your device's memory clear. The more you use your console, the more your RAM (random-access memory) gets filled up, and over time some games and applications can fail to release memory back to your system or leak memory even after you've closed them. These leaks and memory hogging eventually slow down your gaming handheld, and worse, it's not something you can prevent. But by simply restarting your gaming handheld regularly, you can keep those effects from building up over time.\nAnother unsung benefit of restarting your gaming handheld is that it allows your system to apply updates effectively. A good ol' restart allows your device to properly integrate all the important stuff packaged into an update (like bug fixes, security patches and performance improvements) into your system. So, by regularly restarting your gaming handheld, you ensure that your device is always up to date. Regular restarts can also help you resolve software glitches before they cause your gaming handheld to freeze or crash.\nYes, hibernating and powering off helps too\nPowering off your handheld gives the same benefits as restarting simply because it's basically the same thing. Meanwhile, hibernating your gaming handheld doesn't give the benefits of a restart or power-off; instead, it allows you to save your exact session to storage when you're not using the device and then pick up where you left off, all with little to no battery drain. Powering off or hibernating your gaming handheld when it's not in use is widely recommended, especially when you're carrying it around inside a carry case or a bag with poor ventilation where it risks overheating.\nIn case you're wondering, regularly restarting your gaming handheld doesn't mean you need to do it every few hours. There's no fixed or universally accepted number on how often you should restart your gaming handheld, but a good place to start is once every few days or at least once a week.","published":"Sat, 12 Sep 2026 23:30:00 +0000","author":"staff@engadget.com (Isaac Egbon)","guid":"https://www.engadget.com/2252809/benefits-restarting-gaming-handheld-more-often/","created_at":"2026-09-13T01:36:33.414559","last_synchronized":"2026-09-13T01:36:33.414559","sentiment":{"sentiment":"Positive","score":0.4588,"details":{"neg":0.0,"neu":0.842,"pos":0.158,"compound":0.4588}}},{"feed_name":"Engadget","feed_url":"https://www.engadget.com/rss.xml","title":"One problem with Android Auto can be fixed with a simple update","link":"https://www.engadget.com/2252806/android-auto-problem-fixed-simple-firmware-update/","description":"If you've already tried updating your phone, the car's software might be the issue.","content":"One problem with Android Auto can be fixed with a simple update\nYour phone isn't the only device that may need an update when Android Auto starts acting up.\nThere are many reasons why Android Auto may stop connecting, and the initial fixes typically focus on the phone and connection. You inspect the cable, restart the device, forget the car, reconnect it and make sure Android Auto is up to date. If all that fails, however, the software that requires attention might be in the dashboard.\nYour car's infotainment system also has software, and if it's out of date, it can cause Android Auto connection problems. Google advises users to restart the infotainment system and, if the receiver is an aftermarket Pioneer or Kenwood unit, to visit the manufacturer's website and check for a firmware update.\nAndroid Auto relies on the phone and car working correctly, and if you just focus on the phone, you may be missing half the equation. If none of the above has worked, including changing cables, reconnecting the phone or restarting the phone, then the car's software is another area to consider.\nYour car might need an update as well\nAutomakers have been sending out infotainment updates that can improve Android Auto, and Hyundai's February 2026 infotainment update for ccNC-equipped cars in Europe included improved wireless Android Auto and Apple CarPlay connectivity, quicker reconnection if the phone goes unresponsive and automatic restoration after a restart. Chevrolet makes the point more directly: on applicable vehicles, it lists updating both the phone and vehicle software as the first troubleshooting step when Android Auto won't connect.\nHow you install an update depends on the car. Some systems can download software via Wi-Fi, while others use a USB drive or SD card, and some manufacturers also offer dealer assistance. For example, Uconnect allows supported systems to be connected to Wi-Fi via the vehicle's settings menu.\nTherefore, there is no universal button for updating infotainment firmware. Start at your car manufacturer's website and follow the directions for that infotainment system. If the site asks for your VIN or model details, provide them. If the process is not clear (or the manufacturer states that it must be installed by a dealer), a dealer is the safer route than getting an update file from elsewhere.\nA firmware update won't fix every Android Auto problem\nIf Android Auto stopped working right after a phone or Android Auto update, check Google's Android Auto Known Issues page before changing the car's software. Google recently addressed connection problems affecting Pixel 10 and Samsung S26 phones and was still investigating an S25 connection issue. That timing can help you decide which side of the connection deserves attention first.\nIf Android Auto has stopped working right after a phone or Android Auto update, begin there before making any changes in the car. Google advises users to verify that the car is compatible, check that Android Auto is turned on in the infotainment system and test the USB connection. For wired Android Auto, its current guidance is to use a quality cable that is under 3 feet long.\nUpdating the car is a troubleshooting measure and not a magic bullet. If Android Auto was working before, the standard phone-side troubleshooting has failed and newer infotainment software is available, an update is one more thing to check off before scheduling a service appointment.\nFirmware is not without its restrictions. It's not always able to add Android Auto to hardware that didn't support it to begin with. Uconnect says its systems can't simply be upgraded to add Android Auto or CarPlay, while Uconnect 5 and select Uconnect 4 and 4C systems include both.","published":"Sat, 12 Sep 2026 23:00:00 +0000","author":"staff@engadget.com (Husain Parvez)","guid":"https://www.engadget.com/2252806/android-auto-problem-fixed-simple-firmware-update/","created_at":"2026-09-13T01:36:33.414559","last_synchronized":"2026-09-13T01:36:33.414559","sentiment":{"sentiment":"Neutral","score":0.0,"details":{"neg":0.0,"neu":1.0,"pos":0.0,"compound":0.0}}},{"feed_name":"TechRadar","feed_url":"https://www.techradar.com/rss","title":"NYT Strands hints and answers for Sunday, September 13 (game #924)","link":"https://www.techradar.com/computing/websites-apps/nyt-strands-today-answers-hints-13-september-2026","description":"Looking for NYT Strands answers and hints? Here's all you need to know to solve today's game, including the spangram.","content":"NYT Strands hints and answers for Sunday, September 13 (game #924)\nMy clues will help you solve the NYT's Strands today and keep that streak going\nA new NYT Strands puzzle appears at midnight each day for your time zone – which means that some people are always playing 'today's game' while others are playing 'yesterday's'. If you're looking for Saturday's puzzle instead then click here: NYT Strands hints and answers for Saturday, September 12 (game #923).\nStrands is the NYT's latest word game after the likes of Wordle, Spelling Bee and Connections – and it's great fun. It can be difficult, though, so read on for my Strands hints.\nWant more word-based fun? Then check out my NYT Connections today and Quordle today pages for hints and answers for those games, and Marc's Wordle today page for the original viral word game.\nSPOILER WARNING: Information about NYT Strands today is below, so don't read on if you don't want to know the answers.\nNYT Strands today (game #924) - hint #1 - today's theme\nWhat is the theme of today's NYT Strands?\n• Today's NYT Strands theme is… Top secret\nNYT Strands today (game #924) - hint #2 - clue words\nPlay any of these words to unlock the in-game hints system.\n- TEARS\n- STEAR\n- VAST\n- CRATE\n- CRAVEN\n- GRANTED\n- STOLE\nNYT Strands today (game #924) - hint #3 - spangram letters\nHow many letters are in today's spangram?\n• Spangram has 10 letters\nNYT Strands today (game #924) - hint #4 - spangram position\nWhat are two sides of the board that today's spangram touches?\n• First side: left, 4th row\n• Last side: right, 4th row\nRight, the answers are below, so DO NOT SCROLL ANY FURTHER IF YOU DON'T WANT TO SEE THEM.\nNYT Strands today (game #924) - the answers\nThe answers to today's Strands, game #924, are…\n- MOLE\n- AGENT\n- ASSET\n- PLANT\n- COUNTERSPY\n- OPERATIVE\n- SPANGRAM: UNDERCOVER\n- My rating: Hard\n- My score: 2 hints\nThe theme was an obvious one, but I still found it difficult to locate any game words.\nSign up for breaking news, reviews, opinion, top tech deals, and more.\nFortunately, finding non-game words was less tricky, so I took a couple of hints to get going.\nThis left me with three long words, including a spangram that I really should have seen from the start. I’d like to think that this was a deliberately well hidden set of words but human error is far more likely an explanation for my performance.\nYesterday's NYT Strands answers (Saturday, September 12, game #923)\n- LIBERATE\n- FREE\n- RELEASE\n- DETACH\n- UNLOOSE\n- SPANGRAM: LETITGO\nWhat is NYT Strands?\nStrands is the NYT's not-so-new-any-more word game, following Wordle and Connections. It's now a fully fledged member of the NYT's games stable that has been running for a year and which can be played on the NYT Games site on desktop or mobile.\nI've got a full guide to how to play NYT Strands, complete with tips for solving it, so check that out if you're struggling to beat it each day.\nJohnny is a freelance pop culture journalist who has been writing about the internet, music, football and famous people since the iPhone was just a twinkle in Steve Jobs' eye. Previously known by the pseudonym the Pop Detective, his journalistic career began making up stories about Madonna's addiction to sausage rolls (this is not true by the way). A man of few talents, his career is rich and various and includes the highs of interviewing Elton John and Blur; and the lows of interviewing Right Said Fred, appearing on a Channel 5 documentary about Peter Kay, and fact-checking the instruction manual for a German cooker. Somehow still affording to live in North London he is at his happiest riding his bicycle and shouting at pigeons.\n- Marc McLaren Global Editor in Chief\nYou must confirm your public display name before commenting\nPlease logout and then login again, you will then be prompted to enter your display name.","published":"Sat, 12 Sep 2026 23:00:00 +0000","author":"Johnny Dee","guid":"Qi6Lt47bfNkhKGX755NiyQ","created_at":"2026-09-13T01:36:39.726529","last_synchronized":"2026-09-13T01:36:39.726529","sentiment":{"sentiment":"Positive","score":0.2023,"details":{"neg":0.0,"neu":0.913,"pos":0.087,"compound":0.2023}}},{"feed_name":"TechRadar","feed_url":"https://www.techradar.com/rss","title":"Quordle hints and answers for Sunday, September 13 (game #1693)","link":"https://www.techradar.com/computing/websites-apps/quordle-today-answers-clues-13-september-2026","description":"Looking for Quordle clues? We can help. Plus get the answers to Quordle today and past solutions.","content":"Quordle hints and answers for Sunday, September 13 (game #1693)\nMy clues will help you solve Quordle today and keep that streak going\nA new Quordle puzzle appears at midnight each day for your time zone – which means that some people are always playing 'today's game' while others are playing 'yesterday's'. If you're looking for Saturday's puzzle instead then click here: Quordle hints and answers for Saturday, September 12 (game #1692).\nQuordle was one of the original Wordle alternatives and is still going strong now more than 1,500 games later. It offers a genuine challenge, though, so read on if you need some Quordle hints today — or scroll down further for the answers.\nEnjoy playing word games? You can also check out my NYT Connections today and NYT Strands today pages for hints and answers for those puzzles, while Marc's Wordle today column covers the original viral word game.\nSPOILER WARNING: Information about Quordle today is below, so don't read on if you don't want to know the answers.\nQuordle today (game #1693) — hint #1 — Vowels\nHow many different vowels are in Quordle today?\n• The number of different vowels in Quordle today is 4*.\n* Note that by vowel we mean the five standard vowels (A, E, I, O, U), not Y (which is sometimes counted as a vowel too).\nQuordle today (game #1693) — hint #2 — repeated letters\nDo any of today's Quordle answers contain repeated letters?\n• The number of Quordle answers containing a repeated letter today is 0.\nQuordle today (game #1693) — hint #3 — uncommon letters\nDo the letters Q, Z, X or J appear in Quordle today?\n• No. None of Q, Z, X or J appear among today's Quordle answers.\nQuordle today (game #1693) — hint #4 — starting letters (1)\nDo any of today's Quordle puzzles start with the same letter?\n• The number of today's Quordle answers starting with the same letter is 0.\nIf you just want to know the answers at this stage, simply scroll down. If you're not ready yet then here's one more clue to make things a lot easier:\nQuordle today (game #1693) — hint #5 — starting letters (2)\nWhat letters do today's Quordle answers start with?\n• L\n• S\n• M\n• E\nRight, the answers are below, so DO NOT SCROLL ANY FURTHER IF YOU DON'T WANT TO SEE THEM.\nQuordle today (game #1693) — the answers\nThe answers to today's Quordle, game #1693, are…\nSign up for breaking news, reviews, opinion, top tech deals, and more.\n- LUCID\n- SPRAY\n- MANGE\nDespite playing Quordle hundreds of times I’m still making silly mistakes — like today guessing “anime” when the letter 'I' had already been ruled out.\nSometimes I just get a word in my head and have to put it down.\nOh well, at least it wasn’t costly as I soon came to my senses and guessed MANGE.\nDaily Sequence today (game #1693) — the answers\nThe answers to today's Quordle Daily Sequence, game #1693, are…\n- THEME\n- CAMEL\n- AWARD\n- PROSE\nQuordle answers: The past 20\n- Quordle #1692, Saturday, 12 September: SUPER, BRUSH, RESET, SOWER\n- Quordle #1691, Friday, 11 September: PLUSH, PAINT, RIVAL, AFOUL\n- Quordle #1690, Thursday, 10 September: AGENT, EXILE, TITAN, PUREE\n- Quordle #1689, Wednesday, 9 September: HAPPY, GONER, SMACK, MOTIF\n- Quordle #1688, Tuesday, 8 September: FLUNG, DIARY, HALVE, BERTH\n- Quordle #1687, Monday, 7 September: VISIT, CLANK, FEVER, OCEAN\n- Quordle #1686, Sunday, 6 September: PILOT, HONOR, BOWEL, TAUNT\n- Quordle #1685, Saturday, 5 September: GNASH, GEEKY, WHELP, PUTTY\n- Quordle #1684, Friday, 4 September: QUICK, BEEFY, TRUCK, SNAIL\n- Quordle #1683, Thursday, 3 September: SIGMA, ESSAY, MERRY, LLAMA\n- Quordle #1682, Wednesday, 2 September: INCUR, FELLA, TITAN, FROZE\n- Quordle #1681, Tuesday, 1 September: NASAL, ROUTE, ANKLE, MIGHT\n- Quordle #1680, Monday, 31 August: ALIVE, LOUSE, ABOVE, MOSSY\n- Quordle #1679, Sunday, 30 August: USING, SORRY, FILET, PIPER\n- Quordle #1678, Saturday, 29 August: OVINE, ALLOY, VOICE, CARVE\n- Quordle #1677, Friday, 28 August: WAIST, EDICT, BONUS, DEUCE\n- Quordle #1676, Thursday, 27 August: SHEEN, TRICE, WAGON, BEVEL\n- Quordle #1675, Wednesday, 26 August: OPIUM, EPOCH, BESET, FINER\n- Quordle #1674, Tuesday, 25 August: BLESS, FELON, PORCH, CHOSE\n- Quordle #1673, Monday, 24 August: MAJOR, VERGE, CRUSH, SCOFF\nJohnny is a freelance pop culture journalist who has been writing about the internet, music, football and famous people since the iPhone was just a twinkle in Steve Jobs' eye. Previously known by the pseudonym the Pop Detective, his journalistic career began making up stories about Madonna's addiction to sausage rolls (this is not true by the way). A man of few talents, his career is rich and various and includes the highs of interviewing Elton John and Blur; and the lows of interviewing Right Said Fred, appearing on a Channel 5 documentary about Peter Kay, and fact-checking the instruction manual for a German cooker. Somehow still affording to live in North London he is at his happiest riding his bicycle and shouting at pigeons.\n- Marc McLaren Global Editor in Chief\nYou must confirm your public display name before commenting\nPlease logout and then login again, you will then be prompted to enter your display name.","published":"Sat, 12 Sep 2026 23:00:00 +0000","author":"Johnny Dee","guid":"uGjMZGvcU7dybUC2wRZZ6J","created_at":"2026-09-13T01:36:39.726529","last_synchronized":"2026-09-13T01:36:39.726529","sentiment":{"sentiment":"Positive","score":0.5267,"details":{"neg":0.0,"neu":0.773,"pos":0.227,"compound":0.5267}}},{"feed_name":"TechRadar","feed_url":"https://www.techradar.com/rss","title":"NYT Connections hints and answers for Sunday, September 13 (game #1190)","link":"https://www.techradar.com/gaming/nyt-connections-today-answers-hints-13-september-2026","description":"Looking for NYT Connections answers and hints? Here's all you need to know to solve today's game, plus my commentary on the puzzles.","content":"NYT Connections hints and answers for Sunday, September 13 (game #1190)\nMy clues will help you solve the NYT's Connections puzzle today and keep that streak going\nA new NYT Connections puzzle appears at midnight each day for your time zone – which means that some people are always playing 'today's game' while others are playing 'yesterday's'. If you're looking for Saturday's puzzle instead then click here: NYT Connections hints and answers for Saturday, September 12 (game #1189).\nGood morning! Let's play Connections, the NYT's clever word game that challenges you to group answers in various categories. It can be tough, so read on if you need Connections hints.\nWhat should you do once you've finished? Why, play some more word games of course. I've also got daily Strands hints and answers and Quordle hints and answers articles if you need help for those too, while Marc's Wordle today page covers the original viral word game.\nSPOILER WARNING: Information about NYT Connections today is below, so don't read on if you don't want to know the answers.\nNYT Connections today (game #1190) - today's words\nToday's NYT Connections words are…\n- PORRIDGE\n- GARFIELD\n- DROOPY\n- EMPTY\n- HEATHCLIFF\n- EXCESS\n- GRANT\n- SLACK\n- ODIE\n- PARAMOUNT\n- SAGGING\n- MADISON\n- FLACCID\n- KEWPIE\n- CLEVELAND\n- DOUBLESPEAK\nNYT Connections today (game #1190) - hint #1 - group hints\nWhat are some clues for today's NYT Connections groups?\n- YELLOW: Squishy\n- GREEN: American leaders\n- BLUE: Double characters\n- PURPLE: Geographical endings\nNeed more clues?\nWe're firmly in spoiler territory now, but read on if you want to know what the four theme answers are for today's NYT Connections puzzles…\nSign up for breaking news, reviews, opinion, top tech deals, and more.\nNYT Connections today (game #1190) - hint #2 - group answers\nWhat are the answers for today's NYT Connections groups?\n- YELLOW: LACKING IN FIRMNESS\n- GREEN: U.S. PRESIDENTS\n- BLUE: WORDS THAT SOUND LIKE TWO LETTERS\n- PURPLE: ENDING IN ELEVATED LANDFORMS\nRight, the answers are below, so DO NOT SCROLL ANY FURTHER IF YOU DON'T WANT TO SEE THEM.\nNYT Connections today (game #1190) - the answers\nThe answers to today's Connections, game #1190, are…\n- YELLOW: LACKING IN FIRMNESS: DROOPY, FLACCID, SAGGING, SLACK\n- GREEN: U.S. PRESIDENTS: CLEVELAND, GARFIELD, GRANT, MADISON\n- BLUE: WORDS THAT SOUND LIKE TWO LETTERS: EMPTY, EXCESS, KEWPIE, ODIE\n- PURPLE: ENDING IN ELEVATED LANDFORMS: DOUBLESPEAK, HEATHCLIFF, PARAMOUNT, PORRIDGE\n- My rating: Hard\n- My score: 1 mistake\nNot being as knowledgeable about US presidents as I should be, it took me until I had just eight tiles left to connect CLEVELAND, GARFIELD, GRANT, and MADISON.\nEarlier, I had assumed the GARFIELD in question was the cartoon cat and we were looking for a quartet of comic strip felines — HEATHCLIFF is another and I took a punt on DROOPY and CLEVELAND.\nThe silver lining to my error was that it made me notice “cliff” at the end of HEATHCLIFF and I was able to spot the four tiles that made up ENDING IN ELEVATED LANDFORMS, thus snagging a relatively early purple group.\nYesterday's NYT Connections answers (Saturday, September 12, 2026, game #1189)\n- YELLOW: WAYS TO PERSONALIZE A CAR: BUMPER STICKER, FUZZY DICE, HOOD ORNAMENT, VANITY PLATE\n- GREEN: PEOPLE IN AN ONLINE FORUM: LURKER, POSTER, SPAMMER, TROLL\n- BLUE: THINGS THAT BLINK: CURSOR, EYELIDS, SMOKE DETECTOR, TURN SIGNAL\n- PURPLE: ______ ALERT: NERD, RED, SPOILER, WEATHER\nWhat is NYT Connections?\nNYT Connections is one of several increasingly popular word games made by the New York Times. It challenges you to find groups of four items that share something in common, and each group has a different difficulty level: green is easy, yellow a little harder, blue often quite tough and purple usually very difficult.\nOn the plus side, you don't technically need to solve the final one, as you'll be able to answer that one by a process of elimination. What's more, you can make up to four mistakes, which gives you a little bit of breathing room.\nIt's a little more involved than something like Wordle, however, and there are plenty of opportunities for the game to trip you up with tricks. For instance, watch out for homophones and other word games that could disguise the answers.\nIt's playable for free via the NYT Games site on desktop or mobile.\nJohnny is a freelance pop culture journalist who has been writing about the internet, music, football and famous people since the iPhone was just a twinkle in Steve Jobs' eye. Previously known by the pseudonym the Pop Detective, his journalistic career began making up stories about Madonna's addiction to sausage rolls (this is not true by the way). A man of few talents, his career is rich and various and includes the highs of interviewing Elton John and Blur; and the lows of interviewing Right Said Fred, appearing on a Channel 5 documentary about Peter Kay, and fact-checking the instruction manual for a German cooker. Somehow still affording to live in North London he is at his happiest riding his bicycle and shouting at pigeons.\n- Marc McLaren Global Editor in Chief\nYou must confirm your public display name before commenting\nPlease logout and then login again, you will then be prompted to enter your display name.","published":"Sat, 12 Sep 2026 23:00:00 +0000","author":"Johnny Dee","guid":"3GunhiFea6LTaAC6qAVqfM","created_at":"2026-09-13T01:36:39.726529","last_synchronized":"2026-09-13T01:36:39.726529","sentiment":{"sentiment":"Positive","score":0.2023,"details":{"neg":0.0,"neu":0.924,"pos":0.076,"compound":0.2023}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"Ahead of the Chatbot generation: Scaling Production-Ready Agent Fleets with AWS AgentCore","link":"https://dev.to/safiya_k/ahead-of-the-chatbot-generation-scaling-production-ready-agent-fleets-with-aws-agentcore-4a18","description":"<p>Introduction:<br />\nThe days of standalone, text-only chatbots are now a thing of the past. In their place, Agentic AI—systems that can reason independently, plan across multiple steps, manage dynamic memory, and execute tools with self-correction—have become the new standard. In today’s cloud environment, the emphasis has moved from creating basic conversational interfaces to designing robust, highly reliable automation systems that operate on behalf of users.<br />\nWhile the broader technology sector spent months developing unreliable AI agent prototypes using open-source scripts and unstable local frameworks, Amazon Web Services (AWS) took an absolute distinct approach.<br />\nIt completely rebuilt the underlying infrastructure from scratch. With the general availability of Amazon Bedrock AgentCore and the launch of the centralized AWS Agent Registry, AWS has impressively changed the conversation. Now, for technology leaders, the central question is not \"how do we build a single agent?\" but rather \"how do we manage, secure, monitor, and scale an enterprise-wide fleet of agents?\"</p>\n\n<p>For AWS Community Builders, solutions architects, and technology executives, this operational shift represents a significant milestone.<br />\nTransitioning from isolated experiments to production-grade automation means moving beyond conventional software practices. This article provides a comprehensive overview of the architectural challenges involved in safely scaling autonomous agents on AWS infrastructure, using a real-world industry framework to illustrate these advanced capabilities under strict enterprise conditions.</p>\n\n<p>The Core Problem: <br />\nThe Fragility of \"Shadow AI\"</p>\n\n<p>Creating a basic AI agent that checks the weather, drafts an email, or queries a single database table can be done in under an hour using modern APIs.<br />\nHowever, moving such an agent into a highly regulated enterprise setting introduces three major challenges that traditional application monitoring and logging tools cannot address:<br />\n1.<br />\nSilent Failures and Fabricated Tool Execution<br />\nThe most dangerous agent failures are not those that result in explicit error messages, system crashes, or standard HTTP 500 errors.<br />\nInstead, they are silent failures. For example, if an LLM-driven agent fails to understand a complex database schema or encounters an unhandled API timeout, it often generates a confident, realistic, but entirely incorrect response instead of halting execution. In financial, medical, or supply chain workflows, these silent failures pose serious operational risks, leading to corrupted data and poor business decisions without triggering any system alerts.<br />\n2.<br />\nThe Rise of \"Shadow Agents\"<br />\nAs engineering teams rapidly integrate AI capabilities into internal applications, standard cloud governance practices often fail.<br />\nIndependent development teams deploy unmonitored agents across different AWS accounts, using varying foundation models, prompt techniques, and hardcoded API keys.This results in an uncontrolled network of \"Shadow AI\" that bypasses corporate compliance, data loss prevention (DLP) measures, and cost tracking, exposing the enterprise to security threats and uncontrolled cloud spending.<br />\n3.<br />\nContext Collapse in Long-Running Transactions<br />\nStateless APIs struggle with multi-step business processes.<br />\nWhen an autonomous agent is tasked with a complex process—such as processing an insurance claim, verifying documents across three legacy systems, and granting approval—the transaction may take hours or even days.Without a dedicated state management system and long-term memory runtime, agents face \"context collapse,\" which can lead to losing track of their main task, entering infinite loops, or dropping key variables during the process.</p>\n\n\n\n\n<p>The Solution:<br />\nThe AWS Production Agent Stack <br />\nAWS tackles these production vulnerabilities by separating the underlying model, or \"brain,\" from the layers responsible for orchestration, security, and tracking.<br />\nInstead of requiring developers to embed complex state-machine logic and security parameters directly into the foundation model prompt, the modern AWS agent ecosystem abstracts these requirements into three specialized infrastructure layers: the AWS Agent Registry for organization-wide detection and governance, the Bedrock AgentCore Runtime for managing memory, concurrency, state, and tool policies, and the AgentCore Gateway for securing MCP servers and legacy database connectors.</p>\n\n\n\n\n<p>Pillar 1: Decoupled Tool Governance via AgentCore Policies<br />\nTraditionally, if you wanted to restrict what an AI agent could do, you had to hardcode the limitations directly into the LLM prompt (e.g., \"You are not allowed to access table X\") or implement complex conditional logic in your application layer.<br />\nPrompt engineering is inherently unpredictable; advanced prompt injection attacks can easily bypass these restrictions.</p>\n\n<p>With the introduction of Bedrock AgentCore Policies, detailed organizational controls are now fully separated from the agent’s core code.<br />\nThis allows security and compliance teams to create and implement deterministic guardrails that can monitor and intercept tool calls in real time, well before any execution request reaches an external API endpoint.</p>\n\n\n\n\n<p>Pillar 2: Combating \"Shadow AI\" through the AWS Agent Registry<br />\nAs enterprise use of AI expands from just a few agents to hundreds, keeping track of all AI-related resources becomes increasingly difficult for administrators.<br />\nTo address the issue of scattered AI assets across multiple organizations, AWS has introduced Organization-Wide Auto-Detection using the AWS Agent Registry.<br />\nWhen enabled at the root level of AWS Organizations, the registry continuously checks all linked cloud accounts for active Bedrock agent runtimes, Lambda-based tools, and custom model endpoints.<br />\nThese discovered resources are automatically listed on a central \"Detected Endpoints\" dashboard, accessible to IT administrators and compliance officers.This single dashboard allows teams to monitor model usage, token consumption, and overall system latency.</p>\n\n\n\n\n<p>Pillar 3: Simplifying Integration via the Model Context Protocol (MCP)<br />\nIn the past, integration was the most time-consuming part of building agents.<br />\nDevelopers often spent many hours creating custom API wrappers, matching JSON schemas, and handling OAuth credentials for each database, CRM, or internal SaaS platform the agent needed to access.<br />\nTo tackle this integration challenge, AWS has adopted the open-source Model Context Protocol (MCP).<br />\nMCP offers a common, standardized approach that defines how large language models can securely retrieve data and provide tools to external applications.Instead of managing numerous unique API connectors, infrastructure teams can deploy a single MCP server instance that functions as a secure data bridge.<br />\nThrough direct integration with tools like Amazon Quick, business units can find and connect with verified agents without needing to write any code.A business analyst can search the central AWS Agent Registry to locate pre-built MCP servers.With just a few simple clicks, they can securely link an enterprise agent directly to a production data warehouse.</p>\n\n<p>Real-World Industry Implementation: Financial Fraud &amp; Credit Underwriting<br />\nTo better understand the capabilities of this architecture, let's look at a real-world example in the financial services industry: An Automated Fraud Remediation and Credit Underwriting Fleet.</p>\n\n<p>In high-stakes banking environments, an agent fleet must connect with core banking databases, external credit bureaus, customer verification systems, and transaction ledgers.<br />\nThis setting demands multi-turn reasoning, integration with legacy systems, and strict human-in-the-loop triggers to meet financial regulations.<br />\nThe Financial Agent Architecture In Action<br />\nImagine a customer reports an unauthorized company charge.<br />\nA dedicated Fraud Discovery Agent is created to manage the remediation process. The agent must perform the following multi-step workflow across different corporate systems:<br />\n1.<br />\nQuery Transaction Ledgers: Access historical card transactions to confirm the disputed amount.<br />\n2.<br />\nPull Credit Bureau Metrics: Check the merchant's risk profile using external API calls.<br />\n3.<br />\nTrigger Customer Authentication: Send a secure push notification via Amazon Connect to verify the cardholder's identity.<br />\n4.<br />\nIssue Temporary Credit: Return funds to the user account if the transaction meets compliance standards.<br />\nWithout AgentCore, this workflow would rely on hard-coded scripts that could easily cause incorrect fund transfers if the LLM misinterpreted a prompt parameter.</p>\n\n<p>Applying Bedrock AgentCore Policies<br />\nIn the AgentCore architecture, the bank’s risk compliance team sets up an infrastructure-level policy separate from the developer’s application code.<br />\nThe policy states that any automated fund transfer over $500.00 must be stopped immediately and sent to a human manager.<br />\nWhen the agent assesses the fraud claim and tries to execute a tool call for an adjustment of $1,200.00, the AgentCore Policy engine captures the request.<br />\nThe model text does not need to handle this exception. The infrastructure detects the boundary violation, stops the agent's process, and triggers an alert through an internal Amazon Simple Notification Service (SNS) topic to the operations dashboard. The agent is put into a paused state until a human manager reviews the case and confirms the action, combining AI productivity with corporate safeguards.<br />\nArchitecture Spotlight: <br />\nReal-Time Verification Traces<br />\nA key characteristic of a production-ready agent is its ability to be observed.<br />\nIn an enterprise setting, black-box systems are not acceptable. If an agent performs an incorrect transaction, engineers must be able to review the exact sequence of thoughts, tools, and variables that led to the outcome.<br />\nWhen used with Amazon OpenSearch Service MCP Apps, operations teams can access real-time verification traces.<br />\nEvery step in the perception process, from the initial user request to the final API response, is completely visible. If an agent encounters an issue or hits an unexpected policy restriction, the system smoothly transitions from an infrastructure alert to an inline log trace. Given this data directly maps to the developer’s local integrated development environment (IDE), debugging autonomous workflows is as simple as troubleshooting a standard microservice.<br />\nOperational Blueprint: Moving from Prototype to Production<br />\nFor technology executives and cloud architects designing an implementation plan, moving to a production-ready agent fleet requires clear, structured steps:<br />\n• 1.<br />\nCentralize: Identify and list existing \"Shadow AI\" across all departments using the AWS Agent Registry to achieve full visibility.<br />\n• 2.<br />\nDecouple: Remove hardcoded validation rules from application code and move them to deterministic Bedrock AgentCore Policies to ensure compliance.<br />\n• 3.<br />\nStandardize: Convert all custom API integrations to the Model Context Protocol using MCP Servers and Amazon Quick to remove integration challenges.<br />\n• 4.<br />\nObserve: Send execution traces to operational dashboards via OpenSearch Service MCP Apps to maintain full audit trails for compliance and debugging.<br />\nEngineering Deep Dive: Tool Integration with MCP<br />\nTo help cloud engineers understand how tools are made accessible to the Bedrock runtime without manual coding, it’s useful to see how tools are exposed to the Bedrock runtime without writing custom pipelines.<br />\nInstead of creating custom parsing scripts, developers use standard schemas through the Model Context Protocol.</p>\n\n<p>When setting up a tool such as credit issuance, the parameters specify the target account's alphanumeric identifier and a positive decimal value in US dollars.</p>\n\n<p>When this structured configuration is uploaded to the AWS Agent Registry, Bedrock AgentCore automatically reads the operational parameters.<br />\nIt presents this clear data format directly to the foundation model, ensuring the large language model (LLM) organizes its internal reasoning process into the exact structure required by the banking API.This significantly reduces structural errors and inconsistencies in the schema.</p>\n\n\n\n\n<p>Enterprise Governance and Compliance (SOC2 &amp; HIPAA)<br />\nImplementing autonomous agent systems in sectors like finance or healthcare demands strict adherence to standards such as SOC2 Type II, HIPAA, and PCI-DSS.<br />\nThe AWS Agent stack is designed to make these compliance checks easier.<br />\nData Isolation and Cryptographic Enclaves<br />\nAll activities within Bedrock AgentCore ensure complete data isolation.<br />\nBy default, prompt logs, execution paths, and memory contexts are entirely confined within the customer’s virtual private cloud (VPC) environment.AWS does not use enterprise execution data to train public base models.In addition, all memory caches used by long-running agents are encrypted at rest using customer-managed keys through AWS Key Management Service (KMS), meeting strict corporate data protection requirements.<br />\nImmutable Audit Logging for System Trials<br />\nIf a compliance officer or external auditor requests a system review, security teams can retrieve immutable logs from the centralized AWS Agent Registry.<br />\nEvery tool call, model execution, policy restriction, and human approval is recorded with a timestamp and securely stored in Amazon S3 buckets with Object Lock.This level of transparency turns autonomous agents from a potentially risky experiment into a controlled institutional resource that meets strict regulatory requirements.</p>\n\n<p>Conclusion:<br />\nThe Operational Playbook for Enterprise Scale<br />\nCreating an AI agent is no longer simply about solving algorithmic or machine learning data challenges; it has evolved into a matter of software operations and governance.<br />\nThe organizations that are achieving tangible results from agentic workflows are those that are moving away from weak, single-agent proofs of concept toward strong, managed cloud environments.</p>\n\n<p>By dwelling on the foundation of Amazon Bedrock AgentCore, applying clear AgentCore Policies, and maintaining global oversight through the AWS Agent Registry, you can ensure that your autonomous agent teams stay secure, transparent, and fully in line with your core business goals.<br />\nAs we continue to advance into the era of autonomous software systems, the architecture you illustrate today will significantly influence the operational efficiency of your enterprise in the future.</p>","content":"Introduction:\nThe days of standalone, text-only chatbots are now a thing of the past. In their place, Agentic AI—systems that can reason independently, plan across multiple steps, manage dynamic memory, and execute tools with self-correction—have become the new standard. In today’s cloud environment, the emphasis has moved from creating basic conversational interfaces to designing robust, highly reliable automation systems that operate on behalf of users.\nWhile the broader technology sector spent months developing unreliable AI agent prototypes using open-source scripts and unstable local frameworks, Amazon Web Services (AWS) took an absolute distinct approach.\nIt completely rebuilt the underlying infrastructure from scratch. With the general availability of Amazon Bedrock AgentCore and the launch of the centralized AWS Agent Registry, AWS has impressively changed the conversation. Now, for technology leaders, the central question is not \"how do we build a single agent?\" but rather \"how do we manage, secure, monitor, and scale an enterprise-wide fleet of agents?\"\nFor AWS Community Builders, solutions architects, and technology executives, this operational shift represents a significant milestone.\nTransitioning from isolated experiments to production-grade automation means moving beyond conventional software practices. This article provides a comprehensive overview of the architectural challenges involved in safely scaling autonomous agents on AWS infrastructure, using a real-world industry framework to illustrate these advanced capabilities under strict enterprise conditions.\nThe Core Problem:\nThe Fragility of \"Shadow AI\"\nCreating a basic AI agent that checks the weather, drafts an email, or queries a single database table can be done in under an hour using modern APIs.\nHowever, moving such an agent into a highly regulated enterprise setting introduces three major challenges that traditional application monitoring and logging tools cannot address:\n1.\nSilent Failures and Fabricated Tool Execution\nThe most dangerous agent failures are not those that result in explicit error messages, system crashes, or standard HTTP 500 errors.\nInstead, they are silent failures. For example, if an LLM-driven agent fails to understand a complex database schema or encounters an unhandled API timeout, it often generates a confident, realistic, but entirely incorrect response instead of halting execution. In financial, medical, or supply chain workflows, these silent failures pose serious operational risks, leading to corrupted data and poor business decisions without triggering any system alerts.\n2.\nThe Rise of \"Shadow Agents\"\nAs engineering teams rapidly integrate AI capabilities into internal applications, standard cloud governance practices often fail.\nIndependent development teams deploy unmonitored agents across different AWS accounts, using varying foundation models, prompt techniques, and hardcoded API keys.This results in an uncontrolled network of \"Shadow AI\" that bypasses corporate compliance, data loss prevention (DLP) measures, and cost tracking, exposing the enterprise to security threats and uncontrolled cloud spending.\n3.\nContext Collapse in Long-Running Transactions\nStateless APIs struggle with multi-step business processes.\nWhen an autonomous agent is tasked with a complex process—such as processing an insurance claim, verifying documents across three legacy systems, and granting approval—the transaction may take hours or even days.Without a dedicated state management system and long-term memory runtime, agents face \"context collapse,\" which can lead to losing track of their main task, entering infinite loops, or dropping key variables during the process.\nThe Solution:\nThe AWS Production Agent Stack\nAWS tackles these production vulnerabilities by separating the underlying model, or \"brain,\" from the layers responsible for orchestration, security, and tracking.\nInstead of requiring developers to embed complex state-machine logic and security parameters directly into the foundation model prompt, the modern AWS agent ecosystem abstracts these requirements into three specialized infrastructure layers: the AWS Agent Registry for organization-wide detection and governance, the Bedrock AgentCore Runtime for managing memory, concurrency, state, and tool policies, and the AgentCore Gateway for securing MCP servers and legacy database connectors.\nPillar 1: Decoupled Tool Governance via AgentCore Policies\nTraditionally, if you wanted to restrict what an AI agent could do, you had to hardcode the limitations directly into the LLM prompt (e.g., \"You are not allowed to access table X\") or implement complex conditional logic in your application layer.\nPrompt engineering is inherently unpredictable; advanced prompt injection attacks can easily bypass these restrictions.\nWith the introduction of Bedrock AgentCore Policies, detailed organizational controls are now fully separated from the agent’s core code.\nThis allows security and compliance teams to create and implement deterministic guardrails that can monitor and intercept tool calls in real time, well before any execution request reaches an external API endpoint.\nPillar 2: Combating \"Shadow AI\" through the AWS Agent Registry\nAs enterprise use of AI expands from just a few agents to hundreds, keeping track of all AI-related resources becomes increasingly difficult for administrators.\nTo address the issue of scattered AI assets across multiple organizations, AWS has introduced Organization-Wide Auto-Detection using the AWS Agent Registry.\nWhen enabled at the root level of AWS Organizations, the registry continuously checks all linked cloud accounts for active Bedrock agent runtimes, Lambda-based tools, and custom model endpoints.\nThese discovered resources are automatically listed on a central \"Detected Endpoints\" dashboard, accessible to IT administrators and compliance officers.This single dashboard allows teams to monitor model usage, token consumption, and overall system latency.\nPillar 3: Simplifying Integration via the Model Context Protocol (MCP)\nIn the past, integration was the most time-consuming part of building agents.\nDevelopers often spent many hours creating custom API wrappers, matching JSON schemas, and handling OAuth credentials for each database, CRM, or internal SaaS platform the agent needed to access.\nTo tackle this integration challenge, AWS has adopted the open-source Model Context Protocol (MCP).\nMCP offers a common, standardized approach that defines how large language models can securely retrieve data and provide tools to external applications.Instead of managing numerous unique API connectors, infrastructure teams can deploy a single MCP server instance that functions as a secure data bridge.\nThrough direct integration with tools like Amazon Quick, business units can find and connect with verified agents without needing to write any code.A business analyst can search the central AWS Agent Registry to locate pre-built MCP servers.With just a few simple clicks, they can securely link an enterprise agent directly to a production data warehouse.\nReal-World Industry Implementation: Financial Fraud & Credit Underwriting\nTo better understand the capabilities of this architecture, let's look at a real-world example in the financial services industry: An Automated Fraud Remediation and Credit Underwriting Fleet.\nIn high-stakes banking environments, an agent fleet must connect with core banking databases, external credit bureaus, customer verification systems, and transaction ledgers.\nThis setting demands multi-turn reasoning, integration with legacy systems, and strict human-in-the-loop triggers to meet financial regulations.\nThe Financial Agent Architecture In Action\nImagine a customer reports an unauthorized company charge.\nA dedicated Fraud Discovery Agent is created to manage the remediation process. The agent must perform the following multi-step workflow across different corporate systems:\n1.\nQuery Transaction Ledgers: Access historical card transactions to confirm the disputed amount.\n2.\nPull Credit Bureau Metrics: Check the merchant's risk profile using external API calls.\n3.\nTrigger Customer Authentication: Send a secure push notification via Amazon Connect to verify the cardholder's identity.\n4.\nIssue Temporary Credit: Return funds to the user account if the transaction meets compliance standards.\nWithout AgentCore, this workflow would rely on hard-coded scripts that could easily cause incorrect fund transfers if the LLM misinterpreted a prompt parameter.\nApplying Bedrock AgentCore Policies\nIn the AgentCore architecture, the bank’s risk compliance team sets up an infrastructure-level policy separate from the developer’s application code.\nThe policy states that any automated fund transfer over $500.00 must be stopped immediately and sent to a human manager.\nWhen the agent assesses the fraud claim and tries to execute a tool call for an adjustment of $1,200.00, the AgentCore Policy engine captures the request.\nThe model text does not need to handle this exception. The infrastructure detects the boundary violation, stops the agent's process, and triggers an alert through an internal Amazon Simple Notification Service (SNS) topic to the operations dashboard. The agent is put into a paused state until a human manager reviews the case and confirms the action, combining AI productivity with corporate safeguards.\nArchitecture Spotlight:\nReal-Time Verification Traces\nA key characteristic of a production-ready agent is its ability to be observed.\nIn an enterprise setting, black-box systems are not acceptable. If an agent performs an incorrect transaction, engineers must be able to review the exact sequence of thoughts, tools, and variables that led to the outcome.\nWhen used with Amazon OpenSearch Service MCP Apps, operations teams can access real-time verification traces.\nEvery step in the perception process, from the initial user request to the final API response, is completely visible. If an agent encounters an issue or hits an unexpected policy restriction, the system smoothly transitions from an infrastructure alert to an inline log trace. Given this data directly maps to the developer’s local integrated development environment (IDE), debugging autonomous workflows is as simple as troubleshooting a standard microservice.\nOperational Blueprint: Moving from Prototype to Production\nFor technology executives and cloud architects designing an implementation plan, moving to a production-ready agent fleet requires clear, structured steps:\n• 1.\nCentralize: Identify and list existing \"Shadow AI\" across all departments using the AWS Agent Registry to achieve full visibility.\n• 2.\nDecouple: Remove hardcoded validation rules from application code and move them to deterministic Bedrock AgentCore Policies to ensure compliance.\n• 3.\nStandardize: Convert all custom API integrations to the Model Context Protocol using MCP Servers and Amazon Quick to remove integration challenges.\n• 4.\nObserve: Send execution traces to operational dashboards via OpenSearch Service MCP Apps to maintain full audit trails for compliance and debugging.\nEngineering Deep Dive: Tool Integration with MCP\nTo help cloud engineers understand how tools are made accessible to the Bedrock runtime without manual coding, it’s useful to see how tools are exposed to the Bedrock runtime without writing custom pipelines.\nInstead of creating custom parsing scripts, developers use standard schemas through the Model Context Protocol.\nWhen setting up a tool such as credit issuance, the parameters specify the target account's alphanumeric identifier and a positive decimal value in US dollars.\nWhen this structured configuration is uploaded to the AWS Agent Registry, Bedrock AgentCore automatically reads the operational parameters.\nIt presents this clear data format directly to the foundation model, ensuring the large language model (LLM) organizes its internal reasoning process into the exact structure required by the banking API.This significantly reduces structural errors and inconsistencies in the schema.\nEnterprise Governance and Compliance (SOC2 & HIPAA)\nImplementing autonomous agent systems in sectors like finance or healthcare demands strict adherence to standards such as SOC2 Type II, HIPAA, and PCI-DSS.\nThe AWS Agent stack is designed to make these compliance checks easier.\nData Isolation and Cryptographic Enclaves\nAll activities within Bedrock AgentCore ensure complete data isolation.\nBy default, prompt logs, execution paths, and memory contexts are entirely confined within the customer’s virtual private cloud (VPC) environment.AWS does not use enterprise execution data to train public base models.In addition, all memory caches used by long-running agents are encrypted at rest using customer-managed keys through AWS Key Management Service (KMS), meeting strict corporate data protection requirements.\nImmutable Audit Logging for System Trials\nIf a compliance officer or external auditor requests a system review, security teams can retrieve immutable logs from the centralized AWS Agent Registry.\nEvery tool call, model execution, policy restriction, and human approval is recorded with a timestamp and securely stored in Amazon S3 buckets with Object Lock.This level of transparency turns autonomous agents from a potentially risky experiment into a controlled institutional resource that meets strict regulatory requirements.\nConclusion:\nThe Operational Playbook for Enterprise Scale\nCreating an AI agent is no longer simply about solving algorithmic or machine learning data challenges; it has evolved into a matter of software operations and governance.\nThe organizations that are achieving tangible results from agentic workflows are those that are moving away from weak, single-agent proofs of concept toward strong, managed cloud environments.\nBy dwelling on the foundation of Amazon Bedrock AgentCore, applying clear AgentCore Policies, and maintaining global oversight through the AWS Agent Registry, you can ensure that your autonomous agent teams stay secure, transparent, and fully in line with your core business goals.\nAs we continue to advance into the era of autonomous software systems, the architecture you illustrate today will significantly influence the operational efficiency of your enterprise in the future.\nTop comments (0)","published":"Sat, 12 Sep 2026 22:16:28 +0000","author":"Safiya","guid":"https://dev.to/safiya_k/ahead-of-the-chatbot-generation-scaling-production-ready-agent-fleets-with-aws-agentcore-4a18","created_at":"2026-09-13T00:35:43.354162","last_synchronized":"2026-09-13T00:35:43.354162","sentiment":{"sentiment":"Positive","score":0.9958,"details":{"neg":0.064,"neu":0.839,"pos":0.097,"compound":0.9958}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"The Model Wrote the Right Rule and My Replay Rejected It: The Extraction-vs-Replay Split","link":"https://dev.to/debashish_ghosal/the-model-wrote-the-right-rule-and-my-replay-rejected-it-the-extraction-vs-replay-split-4304","description":"<blockquote>\n<p><strong>Update — v0.3.0 released.</strong> CauterRule is now live on <a href=\"https://github.com/deghosal-2026/CauterRule\" rel=\"noopener noreferrer\">GitHub</a> and <a href=\"https://pypi.org/project/cauterule/\" rel=\"noopener noreferrer\">PyPI</a>. It turns repeated agent failures into permanent standing rules — extract, replay-test, promote. <code>pip install cauterule</code> gives you the full CLI, framework adapters, rule lifecycle, pack ecosystem, and official rule packs. The <a href=\"https://github.com/deghosal-2026/CauterRule/blob/main/docs/field-test/v0.3.0/FIELD_TEST_REPORT.md\" rel=\"noopener noreferrer\">v0.3.0 field test report</a> evaluated 2 cloud models across 40 corpora and 4,768 trajectory-runs and is the source for every number below. <a href=\"https://github.com/deghosal-2026/CauterRule/blob/main/docs/release/v0.3.0/release-notes.md\" rel=\"noopener noreferrer\">Release notes</a> · <a href=\"https://github.com/deghosal-2026/CauterRule/blob/main/CHANGELOG.md\" rel=\"noopener noreferrer\">Changelog</a></p>\n</blockquote>\n\n\n\n\n<blockquote>\n<p><strong>CauterRule</strong> is an open-source sidecar that learns standing rules from repeated agent failures. It extracts lessons from trajectories, replay-tests them, and tries to separate reusable guidance from noisy overgeneralization.</p>\n</blockquote>\n\n<p>For two releases we treated \"the pass rate is low\" as one problem and reached for the extractor. It turned out to be two problems we had fused into a single number. This is the split, and it changed where we point next.</p>\n\n<h2>\n\n\nThe assumption: one metric, two questions\n</h2>\n\n<p>The pipeline has two halves that answer two different questions:</p>\n\n<blockquote>\n<p><strong>Extraction:</strong> given a failure, does the model produce the <em>right</em> rule?<br />\n<strong>Replay / evaluation:</strong> given a rule, can we <em>verify</em> it against history?</p>\n</blockquote>\n\n<p>We measured only the second, and we read its score as a verdict on the first. The replay gate that decides promotion computes:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight python\"><code><span class=\"n\">precision</span> <span class=\"o\">=</span> <span class=\"n\">prevented</span> <span class=\"o\">/</span> <span class=\"p\">(</span><span class=\"n\">prevented</span> <span class=\"o\">+</span> <span class=\"n\">broken</span><span class=\"p\">)</span>\n<span class=\"n\">recall</span><span class=\"o\">=</span> <span class=\"n\">prevented</span> <span class=\"o\">/</span> <span class=\"n\">total_failures</span>\n</code></pre>\n\n</div>\n\n\n\n<p>where <code>prevented</code> and <code>broken</code> come from <code>simulate()</code>, which calls <code>rule_matches()</code> — a <strong>text matcher</strong>. So <code>prevented</code> means \"the trigger's <em>prose</em> reached a similarity threshold against a reference failure's prose,\" and <code>broken</code> means \"…against a reference success's prose.\" Nothing in that path asks whether applying the rule's <em>directive</em> would have changed the trajectory's outcome. The gate is a lexical resemblance check wearing a validation costume.</p>\n\n<h2>\n\n\nThe F-001 example: the rule was right, the grader said no\n</h2>\n\n<p>In the <code>failures/positive</code> corpus, trajectory <code>F-001</code> is the canonical git case:</p>\n\n<div class=\"table-wrapper-paragraph\"><table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>\n<code>expected_rule</code> (ground truth)</td>\n<td>\"when git push fails with non-fast-forward, pull latest changes before pushing\"</td>\n</tr>\n<tr>\n<td>extracted <code>when</code>\n</td>\n<td>\"when git push fails with non-fast-forward\"</td>\n</tr>\n<tr>\n<td>extracted <code>do</code>\n</td>\n<td>\"pull latest changes before pushing\"</td>\n</tr>\n</tbody>\n</table></div>\n\n<p>The model reproduced the reference rule almost verbatim. Extraction did its job. Then replay scored it:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight plaintext\"><code>failures_prevented = 5 successes_broken = 3 near_misses = 1\nprecision = 0.625recall = 0.625 verdict = INCONCLUSIVE\n</code></pre>\n\n</div>\n\n\n\n<p>Three \"broken\" successes were <code>S-022-git-pull</code>, <code>S-023-git-status</code>, and <code>NM-044-git-commit-hook</code> — a <em>successful</em> <code>git status</code> counted as broken by a git-push rule, because the two share the token <code>git</code>. A correct rule was demoted to inconclusive by three unrelated successes it merely rhymes with. The model wrote the right rule. The grader rejected it on the prose.</p>\n\n<h2>\n\n\nThe assumption we made (and where it broke)\n</h2>\n\n<blockquote>\n<p><strong>Assumption:</strong> the replay verdict tells us whether the extracted rule is good.</p>\n</blockquote>\n\n<p>It tells us whether the extracted <em>trigger's surface form</em> resembles stored reference <em>surface forms</em>. Those are different claims. A correct rule phrased differently scores low; a wrong rule that happens to share vocabulary scores high. The metric is a <strong>text-similarity proxy</strong>, and it is the gate.</p>\n\n<p>The telling part: the undercount is not unique to replay. The corpus carries a <strong>ground-truth rule</strong> (<code>expected_rule</code>) on 23/50 <code>failures/positive</code> trajectories and 288 <code>reference-expansion</code> trajectories — a ready-made extraction-accuracy instrument we were never scoring against. When we do score it naively (token-F1 of the extracted rule vs <code>expected_rule</code>), we get only ~<strong>0.58</strong> (llama-3.1-8b) / ~<strong>0.50</strong> (gpt-4o-mini). Not because the model is wrong — F-001 is near-verbatim — but because it <em>rewords</em>, and a token comparator can't see through rewording. The same paraphrase gap that breaks replay also undercounts extraction. We had one brittle text-matching lens in two places.</p>\n\n<h2>\n\n\nThe data\n</h2>\n\n<p>Cloud models, <code>failures/positive</code> (23 trajectories carry <code>expected_rule</code>):</p>\n\n<div class=\"table-wrapper-paragraph\"><table>\n<thead>\n<tr>\n<th>Metric</th>\n<th>gpt-4o-mini</th>\n<th>llama-3.1-8b</th>\n<th>What it measures</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Replay pass rate</td>\n<td>8%</td>\n<td>10%</td>\n<td>trigger prose vs reference prose</td>\n</tr>\n<tr>\n<td>Extraction token-F1 vs <code>expected_rule</code>\n</td>\n<td>0.50</td>\n<td>0.58</td>\n<td>extracted rule vs ground truth (text)</td>\n</tr>\n<tr>\n<td>F-001 extracted vs <code>expected_rule</code>\n</td>\n<td>near-verbatim</td>\n<td>near-verbatim</td>\n<td>a correct rule</td>\n</tr>\n<tr>\n<td>F-001 replay verdict</td>\n<td>inconclusive</td>\n<td>inconclusive</td>\n<td>the grader, not the rule</td>\n</tr>\n</tbody>\n</table></div>\n\n<p>The two left-column numbers look bad and point at the model. The right two columns say the model was fine and the <em>measurements</em> are the weak link. The spread is the whole story: extraction is substantively right but lexically variable, and every metric we ship is lexical.</p>\n\n<h2>\n\n\nWhat worked\n</h2>\n\n<ul>\n<li>\n<strong>Naming the two halves separately stopped the misattribution.</strong> \"Low pass rate\" is a symptom; \"extraction accuracy\" and \"replay fidelity\" are the two variables. Once separated, the model stops taking the blame for the grader.</li>\n<li>\n<strong>The ground truth was already in the corpus.</strong> <code>expected_rule</code> needs no new labeling on <code>failures/positive</code> — it was being parsed out and dropped. Picking it up is the cheapest high-signal metric we have.</li>\n<li>\n<strong>One worked example did more than any average.</strong> F-001 — right rule, inconclusive verdict, three token-sharing successes — is the argument in one line. Averages hid it; the single case exposed the mechanism.</li>\n</ul>\n\n<h2>\n\n\nWhat didn't work\n</h2>\n\n<ul>\n<li>\n<strong>We had a ground-truth rule and never scored against it.</strong> <code>expected_rule</code> is dropped at parse (<code>Trajectory</code> has no such field), so the only quality number was the replay verdict. We graded the homework with the wrong rubric for a full release.</li>\n<li>\n<strong>The replay metric masquerades as validation.</strong> \"Tested against history before promotion\" is the product's core promise, but the test is lexical resemblance. A reader who believes the promise and inspects the gate finds prose matching.</li>\n<li>\n<strong>Token-F1 undercounts extraction too.</strong> Reaching only ~0.5–0.6 against known-good rules, a token-F1 extraction score would <em>also</em> wrongly suggest the model is mediocre. The comparator needs a semantic or signature signal, or it repeats the replay's error on the extraction side.</li>\n<li>\n<strong>Both metrics share one brittle primitive.</strong> Because replay and the naive extraction check both lean on the same token matcher, fixing one without the other leaves the split half-measured.</li>\n</ul>\n\n<h2>\n\n\nQuestions we still can't answer\n</h2>\n\n<ul>\n<li>Should replay mean \"does this rule <em>prevent this class</em> of failure\" (behavioral) or \"does it <em>match this history</em>\" (lexical)? We have been shipping the second and calling it the first.</li>\n<li>Can we validate by <em>applying</em> the directive to the reference trajectory and checking the outcome flips — turning <code>prevented</code> from \"text matched\" into \"outcome changed\"?</li>\n<li>Is <code>expected_rule</code> dense enough to backfill onto <code>golden</code> (currently null) so extraction quality has a clean per-corpus number everywhere?</li>\n<li>At what point does a rule that is right-but-reworded deserve credit, and who decides — a threshold, an embedding floor, or a human?</li>\n</ul>\n\n<h2>\n\n\nWhat I learned\n</h2>\n\n<p><strong>A pass rate is a composite; decompose it before you optimize it.</strong> \"8–10% pass\" fused extraction and evaluation. The moment we split the number, the model exonerated itself and the grader took the hit.</p>\n\n<p><strong>Ground truth you don't score against is a decoration.</strong> The corpus already knew the answer (<code>expected_rule</code>). Not wiring it into a metric meant we optimized the wrong half while the right half went unmeasured.</p>\n\n<p><strong>If your \"validation\" only reads words, it can't validate meaning.</strong> The replay gate grades whether the trigger's prose rhymes with stored prose. That is a retrieval property, not a correctness property — and it is the gate.</p>\n\n<p><strong>One crisp failure case beats a dozen averages.</strong> F-001 is the entire diagnosis. When a metric makes a known-good rule fail, the metric is the bug, full stop.</p>\n\n<h2>\n\n\nThe broader lesson\n</h2>\n\n<p>If you build a loop that <em>produces</em> an artifact and then <em>scores</em> it, keep the two measurements separate and keep them honest. A score that grades surface form instead of behavior will make good outputs look bad and bad outputs that share vocabulary look fine — in both halves of the loop at once.</p>\n\n<p>CauterRule spent two releases tuning an extractor that was already writing the right rule, because the number it was chasing was really a prose-similarity score. The fix is not a better model. It is (a) scoring extraction against the <code>expected_rule</code> we already have, with a comparator that can see past rewording, and (b) making the replay gate check whether the <em>directive</em> changes the outcome, not whether the <em>trigger</em> rhymes. Same model. Same corpus. Two honest metrics instead of one flattering lie.</p>\n\n<h2>\n\n\nReferences\n</h2>\n\n<ul>\n<li><a href=\"https://github.com/deghosal-2026/CauterRule/blob/main/docs/release/v0.3.0/release-notes.md\" rel=\"noopener noreferrer\">CauterRule v0.3.0 release notes</a></li>\n<li><a href=\"https://github.com/deghosal-2026/CauterRule/blob/main/docs/field-test/v0.3.0/FIELD_TEST_REPORT.md\" rel=\"noopener noreferrer\">v0.3.0 field test report</a></li>\n<li><a href=\"https://github.com/deghosal-2026/CauterRule/blob/main/docs/corpus/format-spec.md\" rel=\"noopener noreferrer\">Corpus format spec (<code>expected_rule</code>, domain labels)</a></li>\n<li><a href=\"https://github.com/deghosal-2026/CauterRule/blob/main/docs/design/prd/07-success-metrics.md\" rel=\"noopener noreferrer\">PRD — success metrics</a></li>\n<li><a href=\"https://github.com/deghosal-2026/CauterRule/blob/main/docs/design/prd/02-architecture.md\" rel=\"noopener noreferrer\">PRD — architecture</a></li>\n<li>\n<a href=\"https://github.com/deghosal-2026/CauterRule/blob/main/docs/USER_GUIDE.md\" rel=\"noopener noreferrer\">User guide</a> · <a href=\"https://github.com/deghosal-2026/CauterRule/blob/main/CHANGELOG.md\" rel=\"noopener noreferrer\">Changelog</a>\n</li>\n</ul>\n\n\n\n\n<blockquote>\n<p><strong>CauterRule v0.3.0 is released.</strong> The replay and extraction data are in the <a href=\"https://github.com/deghosal-2026/CauterRule/blob/main/docs/field-test/v0.3.0/FIELD_TEST_REPORT.md\" rel=\"noopener noreferrer\">field test report</a>. The <a href=\"https://github.com/deghosal-2026/CauterRule\" rel=\"noopener noreferrer\">repo</a> is public. Install with <code>pip install cauterule</code>. <a href=\"https://github.com/deghosal-2026/CauterRule/blob/main/CHANGELOG.md\" rel=\"noopener noreferrer\">Changelog</a> · <a href=\"https://github.com/deghosal-2026/CauterRule/blob/main/docs/release/v0.3.0/release-notes.md\" rel=\"noopener noreferrer\">Release notes</a></p>\n</blockquote>","content":"Update — v0.3.0 released. CauterRule is now live on GitHub and PyPI. It turns repeated agent failures into permanent standing rules — extract, replay-test, promote.\npip install cauterule\ngives you the full CLI, framework adapters, rule lifecycle, pack ecosystem, and official rule packs. The v0.3.0 field test report evaluated 2 cloud models across 40 corpora and 4,768 trajectory-runs and is the source for every number below. Release notes · Changelog\nCauterRule is an open-source sidecar that learns standing rules from repeated agent failures. It extracts lessons from trajectories, replay-tests them, and tries to separate reusable guidance from noisy overgeneralization.\nFor two releases we treated \"the pass rate is low\" as one problem and reached for the extractor. It turned out to be two problems we had fused into a single number. This is the split, and it changed where we point next.\nThe assumption: one metric, two questions\nThe pipeline has two halves that answer two different questions:\nExtraction: given a failure, does the model produce the right rule?\nReplay / evaluation: given a rule, can we verify it against history?\nWe measured only the second, and we read its score as a verdict on the first. The replay gate that decides promotion computes:\nprecision = prevented / (prevented + broken)\nrecall = prevented / total_failures\nwhere prevented\nand broken\ncome from simulate()\n, which calls rule_matches()\n— a text matcher. So prevented\nmeans \"the trigger's prose reached a similarity threshold against a reference failure's prose,\" and broken\nmeans \"…against a reference success's prose.\" Nothing in that path asks whether applying the rule's directive would have changed the trajectory's outcome. The gate is a lexical resemblance check wearing a validation costume.\nThe F-001 example: the rule was right, the grader said no\nIn the failures/positive\ncorpus, trajectory F-001\nis the canonical git case:\nexpected_rule (ground truth) |\n\"when git push fails with non-fast-forward, pull latest changes before pushing\" |\nextracted when\n|\n\"when git push fails with non-fast-forward\" |\nextracted do\n|\n\"pull latest changes before pushing\" |\nThe model reproduced the reference rule almost verbatim. Extraction did its job. Then replay scored it:\nfailures_prevented = 5 successes_broken = 3 near_misses = 1\nprecision = 0.625 recall = 0.625 verdict = INCONCLUSIVE\nThree \"broken\" successes were S-022-git-pull\n, S-023-git-status\n, and NM-044-git-commit-hook\n— a successful git status\ncounted as broken by a git-push rule, because the two share the token git\n. A correct rule was demoted to inconclusive by three unrelated successes it merely rhymes with. The model wrote the right rule. The grader rejected it on the prose.\nThe assumption we made (and where it broke)\nAssumption: the replay verdict tells us whether the extracted rule is good.\nIt tells us whether the extracted trigger's surface form resembles stored reference surface forms. Those are different claims. A correct rule phrased differently scores low; a wrong rule that happens to share vocabulary scores high. The metric is a text-similarity proxy, and it is the gate.\nThe telling part: the undercount is not unique to replay. The corpus carries a ground-truth rule (expected_rule\n) on 23/50 failures/positive\ntrajectories and 288 reference-expansion\ntrajectories — a ready-made extraction-accuracy instrument we were never scoring against. When we do score it naively (token-F1 of the extracted rule vs expected_rule\n), we get only ~0.58 (llama-3.1-8b) / ~0.50 (gpt-4o-mini). Not because the model is wrong — F-001 is near-verbatim — but because it rewords, and a token comparator can't see through rewording. The same paraphrase gap that breaks replay also undercounts extraction. We had one brittle text-matching lens in two places.\nThe data\nCloud models, failures/positive\n(23 trajectories carry expected_rule\n):\n| Metric | gpt-4o-mini | llama-3.1-8b | What it measures |\n|---|---|---|---|\n| Replay pass rate | 8% | 10% | trigger prose vs reference prose |\nExtraction token-F1 vs expected_rule\n|\n0.50 | 0.58 | extracted rule vs ground truth (text) |\nF-001 extracted vs expected_rule\n|\nnear-verbatim | near-verbatim | a correct rule |\n| F-001 replay verdict | inconclusive | inconclusive | the grader, not the rule |\nThe two left-column numbers look bad and point at the model. The right two columns say the model was fine and the measurements are the weak link. The spread is the whole story: extraction is substantively right but lexically variable, and every metric we ship is lexical.\nWhat worked\n- Naming the two halves separately stopped the misattribution. \"Low pass rate\" is a symptom; \"extraction accuracy\" and \"replay fidelity\" are the two variables. Once separated, the model stops taking the blame for the grader.\n-\nThe ground truth was already in the corpus.\nexpected_rule\nneeds no new labeling onfailures/positive\n— it was being parsed out and dropped. Picking it up is the cheapest high-signal metric we have. - One worked example did more than any average. F-001 — right rule, inconclusive verdict, three token-sharing successes — is the argument in one line. Averages hid it; the single case exposed the mechanism.\nWhat didn't work\n-\nWe had a ground-truth rule and never scored against it.\nexpected_rule\nis dropped at parse (Trajectory\nhas no such field), so the only quality number was the replay verdict. We graded the homework with the wrong rubric for a full release. - The replay metric masquerades as validation. \"Tested against history before promotion\" is the product's core promise, but the test is lexical resemblance. A reader who believes the promise and inspects the gate finds prose matching.\n- Token-F1 undercounts extraction too. Reaching only ~0.5–0.6 against known-good rules, a token-F1 extraction score would also wrongly suggest the model is mediocre. The comparator needs a semantic or signature signal, or it repeats the replay's error on the extraction side.\n- Both metrics share one brittle primitive. Because replay and the naive extraction check both lean on the same token matcher, fixing one without the other leaves the split half-measured.\nQuestions we still can't answer\n- Should replay mean \"does this rule prevent this class of failure\" (behavioral) or \"does it match this history\" (lexical)? We have been shipping the second and calling it the first.\n- Can we validate by applying the directive to the reference trajectory and checking the outcome flips — turning\nprevented\nfrom \"text matched\" into \"outcome changed\"? - Is\nexpected_rule\ndense enough to backfill ontogolden\n(currently null) so extraction quality has a clean per-corpus number everywhere? - At what point does a rule that is right-but-reworded deserve credit, and who decides — a threshold, an embedding floor, or a human?\nWhat I learned\nA pass rate is a composite; decompose it before you optimize it. \"8–10% pass\" fused extraction and evaluation. The moment we split the number, the model exonerated itself and the grader took the hit.\nGround truth you don't score against is a decoration. The corpus already knew the answer (expected_rule\n). Not wiring it into a metric meant we optimized the wrong half while the right half went unmeasured.\nIf your \"validation\" only reads words, it can't validate meaning. The replay gate grades whether the trigger's prose rhymes with stored prose. That is a retrieval property, not a correctness property — and it is the gate.\nOne crisp failure case beats a dozen averages. F-001 is the entire diagnosis. When a metric makes a known-good rule fail, the metric is the bug, full stop.\nThe broader lesson\nIf you build a loop that produces an artifact and then scores it, keep the two measurements separate and keep them honest. A score that grades surface form instead of behavior will make good outputs look bad and bad outputs that share vocabulary look fine — in both halves of the loop at once.\nCauterRule spent two releases tuning an extractor that was already writing the right rule, because the number it was chasing was really a prose-similarity score. The fix is not a better model. It is (a) scoring extraction against the expected_rule\nwe already have, with a comparator that can see past rewording, and (b) making the replay gate check whether the directive changes the outcome, not whether the trigger rhymes. Same model. Same corpus. Two honest metrics instead of one flattering lie.\nReferences\n- CauterRule v0.3.0 release notes\n- v0.3.0 field test report\n- Corpus format spec (\nexpected_rule\n, domain labels) - PRD — success metrics\n- PRD — architecture\n- User guide · Changelog\nCauterRule v0.3.0 is released. The replay and extraction data are in the field test report. The repo is public. Install with\npip install cauterule\n. Changelog · Release notes\nTop comments (0)","published":"Sat, 12 Sep 2026 22:04:00 +0000","author":"Debashish Ghosal","guid":"https://dev.to/debashish_ghosal/the-model-wrote-the-right-rule-and-my-replay-rejected-it-the-extraction-vs-replay-split-4304","created_at":"2026-09-13T00:35:43.354162","last_synchronized":"2026-09-13T00:35:43.354162","sentiment":{"sentiment":"Negative","score":-0.9222,"details":{"neg":0.063,"neu":0.878,"pos":0.058,"compound":-0.9222}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"I ran my scanner against 5 real CVEs. It missed 4. Then I reverted my own fix.","link":"https://dev.to/balbaks/i-ran-my-scanner-against-5-real-cves-it-missed-4-then-i-reverted-my-own-fix-4dhk","description":"<p><strong>Why this post is different from the last one</strong></p>\n\n<p>The last write-up in this series announced four tools. This one is about what happened when I stopped writing tests for my own tools and started checking one of them against reality — and about the fix I built, tested, shipped, and then took back out, because it was wrong in a way that only showed up once I looked past the headline number.</p>\n\n<p><strong>The setup: inlet, and the claim it hadn't actually tested</strong></p>\n\n<p><strong>inlet</strong> is a static scanner: point it at a Python codebase, it finds every call site that looks like SQL execution — raw DB API calls, Django's .raw()/.extra(), SQLAlchemy's text() — and classifies each as parameterized, concatenated, or uncertain. Every claim in its original README was backed by 7 hand-written fixtures, each proving one specific classification rule.</p>\n\n<p>That's real, but it's not evidence of anything beyond \"the mechanism works on cases designed to exercise it.\" So I built a real-world evaluation: 15 PyPI packages, split into two groups.</p>\n\n<p><strong>Group A (5 packages):</strong> each with a documented, independently verified historical SQL-injection CVE, with the fix commit or advisory located ahead of time so I could check inlet's output against ground truth, not against inlet's own opinion of itself.</p>\n\n<p><strong>Group B (10 packages):</strong> popular, no known SQLi history — a noise-floor check on how often inlet flags something that isn't actually a risk.</p>\n\n<p>~772K lines of real code, scanned unmodified, in about 10.6 seconds combined.</p>\n\n<p>The result: 0 for 5</p>\n\n<p>Django (CVE-2022-28346), Apache Superset (CVE-2023-49736), Tortoise ORM (CVE-2020-11010), and Airflow's common-sql provider (CVE-2025-30473) were all complete misses. Archery's CVE-2023-30556 was a partial — the exact vulnerable line showed up in inlet's output, but classified uncertain instead of concatenated, because the f-string assignment sat inside a try: block, a name-resolution gap beyond the documented scope wall.</p>\n\n<p>The pattern behind all four full misses was identical: the vulnerable code never literally calls something named .execute(), .raw(), .extra(), or text(). It goes through a framework's own abstraction — hook.get_records(), field.like(), a plain helper function — that eventually reaches SQL execution, several calls away from any name inlet recognizes.</p>\n\n<p>I put that as the headline of EVALUATION.md, not a footnote. A 0/5 result, reported plainly, is worth more than a clean-looking demo — it's the first piece of evidence in this whole project that came from checking against something I didn't design.</p>\n\n<p><strong>The smaller, real bug the same evaluation found</strong></p>\n\n<p>Group B surfaced something fixable: one package's uncertain findings were 67% false positives, from a name collision. peewee's own query builder has an .execute(database) method — same method name as a real DB cursor's .execute(sql), completely different meaning. inlet's name-only matching couldn't tell them apart.</p>\n\n<p><strong>The fix that worked, and then didn't</strong></p>\n\n<p>I built a positive-evidence rule: only treat an .execute()-shaped call as a real DB-idiom candidate if there's actual evidence for it — either the argument is string-shaped, or the receiver chain shows a .cursor() call or a conventional cursor/connection name. Otherwise, exclude it.</p>\n\n<p>It worked, exactly as intended, on peewee: 33 uncertain findings down to 9, a clean diff confirming all 24 removed were the exact false-positive shape, zero true positives lost.</p>\n\n<p>Then I re-ran the other 9 Group B packages, and found the same rule had silently dropped 94 real database call sites — Django's SchemaEditor.execute(), SQLAlchemy's own Engine/Session internals, a dataset helper, SQLModel's super().execute(). All real DB calls, lost for one reason: their receiver was named something generic like self, which the new rule couldn't distinguish from peewee's unrelated Query.execute().</p>\n\n<p><strong>Why I reverted it instead of tuning it further</strong></p>\n\n<p>There's a version of this where I keep iterating the heuristic, trying to find a cleverer rule that keeps the peewee win without the 94-finding cost. I didn't do that, because the actual finding underneath both results is more important than either number:</p>\n\n<p>When the argument isn't string-shaped and the receiver name is generic, there is no way to tell a real DB wrapper from an unrelated same-named method using local syntax alone. self.execute(x) is genuinely, irreducibly ambiguous from where inlet sits. That's not a heuristic to keep tuning — it's the same category of hard limit as the tool's existing cross-function-scope wall.</p>\n\n<p>And there's an asymmetry that matters more than either number: a finding in uncertain is recoverable — a human can look at it and dismiss it. A finding that's silently excluded is not recoverable — it never existed for anyone to see. Trading visible noise for confident silence is a strictly worse failure mode, even when the summary metric (fewer uncertain findings!) looks like an improvement.</p>\n\n<p>So I reverted it. Every .execute()-shaped candidate goes back to being reported, at whatever verdict the classifier can actually support — no silent exclusion, ever. The receiver-evidence detection code is still there, inert, available as a future upgrade-only signal (never a removal signal) if a principled way to use it that way ever turns up.</p>\n\n<p>I wrote the whole thing up as its own section in EVALUATION.md — what broke, why it was reverted, and the actual finding — because a failed attempt with an honest postmortem is a better artifact than either the original bug or a fix that quietly traded one failure mode for a worse one.</p>\n\n<p><strong>Then, a fifth tool: escrow</strong></p>\n\n<p>Separately, I built escrow, which vets a Python package before a real pip install by actually installing and importing it in a sandbox first — built directly on two earlier tools in this series: husk (the hardened sandbox) and witness (the audit-hook behavior reporter). This exists because of slopsquatting: LLMs hallucinate plausible-but-nonexistent package names at meaningful rates, attackers register those exact names, and the next pip install executes whatever they put there — this is already a real, documented attack pattern, not a hypothetical.</p>\n\n<p>Building it surfaced a real limitation in witness's own technique: witness observes behavior by prepending an audit-hook preamble to a script running in one interpreter process. That can't see into pip's own build-backend subprocess — exactly where install-time (setup.py) attacks actually run. escrow's hook ships instead as a real sitecustomize.py, auto-loaded by Python's own site module in every subprocess pip spawns, not just the top-level driver. Found and fixed empirically, including discovering that pip install silently swallows successful build-step subprocess output unless run with --verbose — which would have hidden a caught-and-ignored malicious write from the very report meant to catch it.</p>\n\n<p>escrow keeps its own honest limit stated up front: installing a real package requires network access, so anything malicious that completes fast enough during that window can be detected and reported, but not prevented in real time. That's not a gap to be engineered around in v0.1.0 — it's a fundamental property of vetting something that needs network access to install at all.</p>\n\n<p><strong>The actual pattern across all five tools now</strong></p>\n\n<p>secfix refuses to say \"fixed\" without a fresh trace. husk backs every hardening claim with an adversarial test. witness turns its own blind spot into a loud signal instead of a silent gap. inlet measured itself against real CVEs, got a bad number, and reported it as the headline. And when a fix improved that number by making the tool quietly worse in a different way, it went back out — documented, not buried.</p>\n\n<p>That's the thing I'm actually trying to build a track record of. Not five clever tools. Five tools that keep finding their own mistakes before anyone else has to.</p>\n\n<p>Repos: github.com/balbaks/secfix · github.com/balbaks/husk · github.com/balbaks/witness · github.com/balbaks/inlet · github.com/balbaks/escrow</p>","content":"Why this post is different from the last one\nThe last write-up in this series announced four tools. This one is about what happened when I stopped writing tests for my own tools and started checking one of them against reality — and about the fix I built, tested, shipped, and then took back out, because it was wrong in a way that only showed up once I looked past the headline number.\nThe setup: inlet, and the claim it hadn't actually tested\ninlet is a static scanner: point it at a Python codebase, it finds every call site that looks like SQL execution — raw DB API calls, Django's .raw()/.extra(), SQLAlchemy's text() — and classifies each as parameterized, concatenated, or uncertain. Every claim in its original README was backed by 7 hand-written fixtures, each proving one specific classification rule.\nThat's real, but it's not evidence of anything beyond \"the mechanism works on cases designed to exercise it.\" So I built a real-world evaluation: 15 PyPI packages, split into two groups.\nGroup A (5 packages): each with a documented, independently verified historical SQL-injection CVE, with the fix commit or advisory located ahead of time so I could check inlet's output against ground truth, not against inlet's own opinion of itself.\nGroup B (10 packages): popular, no known SQLi history — a noise-floor check on how often inlet flags something that isn't actually a risk.\n~772K lines of real code, scanned unmodified, in about 10.6 seconds combined.\nThe result: 0 for 5\nDjango (CVE-2022-28346), Apache Superset (CVE-2023-49736), Tortoise ORM (CVE-2020-11010), and Airflow's common-sql provider (CVE-2025-30473) were all complete misses. Archery's CVE-2023-30556 was a partial — the exact vulnerable line showed up in inlet's output, but classified uncertain instead of concatenated, because the f-string assignment sat inside a try: block, a name-resolution gap beyond the documented scope wall.\nThe pattern behind all four full misses was identical: the vulnerable code never literally calls something named .execute(), .raw(), .extra(), or text(). It goes through a framework's own abstraction — hook.get_records(), field.like(), a plain helper function — that eventually reaches SQL execution, several calls away from any name inlet recognizes.\nI put that as the headline of EVALUATION.md, not a footnote. A 0/5 result, reported plainly, is worth more than a clean-looking demo — it's the first piece of evidence in this whole project that came from checking against something I didn't design.\nThe smaller, real bug the same evaluation found\nGroup B surfaced something fixable: one package's uncertain findings were 67% false positives, from a name collision. peewee's own query builder has an .execute(database) method — same method name as a real DB cursor's .execute(sql), completely different meaning. inlet's name-only matching couldn't tell them apart.\nThe fix that worked, and then didn't\nI built a positive-evidence rule: only treat an .execute()-shaped call as a real DB-idiom candidate if there's actual evidence for it — either the argument is string-shaped, or the receiver chain shows a .cursor() call or a conventional cursor/connection name. Otherwise, exclude it.\nIt worked, exactly as intended, on peewee: 33 uncertain findings down to 9, a clean diff confirming all 24 removed were the exact false-positive shape, zero true positives lost.\nThen I re-ran the other 9 Group B packages, and found the same rule had silently dropped 94 real database call sites — Django's SchemaEditor.execute(), SQLAlchemy's own Engine/Session internals, a dataset helper, SQLModel's super().execute(). All real DB calls, lost for one reason: their receiver was named something generic like self, which the new rule couldn't distinguish from peewee's unrelated Query.execute().\nWhy I reverted it instead of tuning it further\nThere's a version of this where I keep iterating the heuristic, trying to find a cleverer rule that keeps the peewee win without the 94-finding cost. I didn't do that, because the actual finding underneath both results is more important than either number:\nWhen the argument isn't string-shaped and the receiver name is generic, there is no way to tell a real DB wrapper from an unrelated same-named method using local syntax alone. self.execute(x) is genuinely, irreducibly ambiguous from where inlet sits. That's not a heuristic to keep tuning — it's the same category of hard limit as the tool's existing cross-function-scope wall.\nAnd there's an asymmetry that matters more than either number: a finding in uncertain is recoverable — a human can look at it and dismiss it. A finding that's silently excluded is not recoverable — it never existed for anyone to see. Trading visible noise for confident silence is a strictly worse failure mode, even when the summary metric (fewer uncertain findings!) looks like an improvement.\nSo I reverted it. Every .execute()-shaped candidate goes back to being reported, at whatever verdict the classifier can actually support — no silent exclusion, ever. The receiver-evidence detection code is still there, inert, available as a future upgrade-only signal (never a removal signal) if a principled way to use it that way ever turns up.\nI wrote the whole thing up as its own section in EVALUATION.md — what broke, why it was reverted, and the actual finding — because a failed attempt with an honest postmortem is a better artifact than either the original bug or a fix that quietly traded one failure mode for a worse one.\nThen, a fifth tool: escrow\nSeparately, I built escrow, which vets a Python package before a real pip install by actually installing and importing it in a sandbox first — built directly on two earlier tools in this series: husk (the hardened sandbox) and witness (the audit-hook behavior reporter). This exists because of slopsquatting: LLMs hallucinate plausible-but-nonexistent package names at meaningful rates, attackers register those exact names, and the next pip install executes whatever they put there — this is already a real, documented attack pattern, not a hypothetical.\nBuilding it surfaced a real limitation in witness's own technique: witness observes behavior by prepending an audit-hook preamble to a script running in one interpreter process. That can't see into pip's own build-backend subprocess — exactly where install-time (setup.py) attacks actually run. escrow's hook ships instead as a real sitecustomize.py, auto-loaded by Python's own site module in every subprocess pip spawns, not just the top-level driver. Found and fixed empirically, including discovering that pip install silently swallows successful build-step subprocess output unless run with --verbose — which would have hidden a caught-and-ignored malicious write from the very report meant to catch it.\nescrow keeps its own honest limit stated up front: installing a real package requires network access, so anything malicious that completes fast enough during that window can be detected and reported, but not prevented in real time. That's not a gap to be engineered around in v0.1.0 — it's a fundamental property of vetting something that needs network access to install at all.\nThe actual pattern across all five tools now\nsecfix refuses to say \"fixed\" without a fresh trace. husk backs every hardening claim with an adversarial test. witness turns its own blind spot into a loud signal instead of a silent gap. inlet measured itself against real CVEs, got a bad number, and reported it as the headline. And when a fix improved that number by making the tool quietly worse in a different way, it went back out — documented, not buried.\nThat's the thing I'm actually trying to build a track record of. Not five clever tools. Five tools that keep finding their own mistakes before anyone else has to.\nRepos: github.com/balbaks/secfix · github.com/balbaks/husk · github.com/balbaks/witness · github.com/balbaks/inlet · github.com/balbaks/escrow\nTop comments (0)","published":"Sat, 12 Sep 2026 21:51:23 +0000","author":"Hassan Balbakie","guid":"https://dev.to/balbaks/i-ran-my-scanner-against-5-real-cves-it-missed-4-then-i-reverted-my-own-fix-4dhk","created_at":"2026-09-13T00:35:43.354162","last_synchronized":"2026-09-13T00:35:43.354162","sentiment":{"sentiment":"Negative","score":-0.97,"details":{"neg":0.093,"neu":0.831,"pos":0.076,"compound":-0.97}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"Transcribing a Lecture and Transcribing a Caller Are Not the Same Problem","link":"https://dev.to/nabeelbaghoor/transcribing-a-lecture-and-transcribing-a-caller-are-not-the-same-problem-45j0","description":"<p>The first time I opened a Retell call transcript, I assumed I already knew this problem.</p>\n\n<p>I had shipped LectureNotes AI before that: a note-taking app that records a lecture and turns it into a summary and a revision outline. Audio in, text out, users happy. So when voice agents became most of my work, I filed speech recognition under solved and moved on to the parts I thought were hard.</p>\n\n<p>I was wrong in a way that cost me a few weeks. Transcribing a fifty minute lecture and transcribing a caller on a phone line share a model family and share almost nothing else. They fail differently, they are evaluated differently, and the thing that improves each one is in a completely different layer of the stack.</p>\n\n<p>Here is the honest split, after shipping both.</p>\n\n<h2>\n\n\nOne is a batch job, one is a turn\n</h2>\n\n<p>Long-form transcription is a batch problem. A student hits record, sits through an hour, closes the laptop. The audio arrives as one file. Nobody is waiting on a socket for the next token. If a pass comes out badly you can run it again with different settings and nobody knows. You can spend three minutes of compute on an hour of audio and the user experiences that as \"it was ready when I looked\".</p>\n\n<p>A voice agent inverts every one of those properties. Audio arrives as a stream, the useful unit is a three to eight second turn, and the transcript is consumed immediately by a model that is about to say something out loud to a human being. There is no second pass and no running it again with a bigger model. The caller is on the line, and every millisecond you spend is silence they are listening to.</p>\n\n<p>Once you frame it that way, the question stops being which STT provider is best and starts being which failure you can afford.</p>\n\n<h2>\n\n\nIn batch, the capture path beats the model\n</h2>\n\n<p>The single biggest accuracy win on the note-taking side had nothing to do with the transcription model. It was recording conditions.</p>\n\n<p>A lecture recording is a phone lying on a desk six rows back, in a room with hard walls and an air conditioner. The lecturer walks around. Against that, the gap between two good STT models is noise.</p>\n\n<p>What actually helped was unglamorous: handling the recording session properly so it survives the screen locking and the app being backgrounded, being honest in the UI about what a bad recording is going to produce, and getting the sample rate and format right at the source instead of repairing it downstream. If you upsample rubbish you get expensive rubbish.</p>\n\n<p>If you are building anything that ingests audio, that is where I would spend the first week, not on provider benchmarks.</p>\n\n<h2>\n\n\nWord error rate measures the wrong words\n</h2>\n\n<p>WER treats every word as equal. Users do not.</p>\n\n<p>A lecture transcript that is broadly correct but mangles the lecturer's specific terminology is useless, because the terminology is the entire reason a student is reading it. Course jargon, people's names, formula names, abbreviations: those are exactly the tokens with the least redundancy in the surrounding sentence, so context cannot repair them. Meanwhile a transcript that quietly drops filler words and tidies false starts usually tests better than a faithful one.</p>\n\n<p>So the useful evaluation was never an aggregate score. It was: did the terms survive? Feeding known domain vocabulary in as a bias list or prompt hint bought more perceived quality than any model swap I tried, because it targeted the small set of tokens carrying all the meaning.</p>\n\n<p>The same shape shows up on the phone side with different tokens. Identifiers have no redundancy either: nothing around a postcode, a registration number or an email address tells the model what it should have been.</p>\n\n<p>Pick the twenty terms or the five fields that actually matter to your user and measure those. An aggregate score will keep telling you things are fine while the only words anyone cares about are wrong.</p>\n\n<h2>\n\n\nWhat happens after the transcript is the actual product\n</h2>\n\n<p>For long-form, the transcript is not the deliverable. The summary and the outline are, and that is where the quality is genuinely won.</p>\n\n<p>Three things mattered more than I expected:</p>\n\n<p><strong>Chunk on the audio, not on the character count.</strong> Cutting a transcript into fixed windows regularly slices a definition in half, and the summary then confidently loses it. Chunking around natural pauses, with overlap between chunks, produced noticeably better output for the same model and the same prompt.</p>\n\n<p><strong>Keep timestamps all the way through.</strong> If a student cannot jump from a summary line back to the moment it came from, they cannot check it, and if they cannot check it they will not trust it. Traceability is a feature, not plumbing.</p>\n\n<p><strong>Structure is a product decision, not a prompt afterthought.</strong> Students did not want prose. They wanted something scannable: headings, short lines, definitions pulled out. An excellent essay-shaped summary tested worse than a mediocre outline-shaped one.</p>\n\n<p>I run the same instinct on voice work now. The transcript is raw signal. Everything the client actually cares about, the booking, the CRM record, the outcome label, is a structured extraction sitting on top of it.</p>\n\n<h2>\n\n\nOn a phone line, you are transcribing through a straw\n</h2>\n\n<p>Then there is the call.</p>\n\n<p>Telephony audio is narrowband. This is the part I wish someone had said to me plainly: it is not a lightly degraded version of good audio, it is audio with the top of the spectrum removed, and the frequencies it removes are the ones that distinguish consonants. That is why f and s collapse into each other, why m and n become a coin toss, and why five and nine are a permanent problem when you are taking a phone number. It is not the model failing. That information never reached the model.</p>\n\n<p>On top of that you get callers in cars, on speaker, standing in a busy shop, with regional accents that a US-tuned recognition setting handles badly.</p>\n\n<p>Because the audio is worse and the stakes are higher, the fix has to leave the STT layer entirely. You cannot make the line better. You can design a conversation that survives mishearing: confirm identifiers in small chunks as you take them, validate shape downstream in automation rather than hoping the prompt caught it, and build a repair ladder with an exit, so a bad capture becomes a text message or a transfer instead of a loop.</p>\n\n<h2>\n\n\nLatency is the product in exactly one of them\n</h2>\n\n<p>In batch, latency is a scheduling detail. In a call, it is the product.</p>\n\n<p>Streaming STT adds cost at the point where you are most sensitive. Endpointing has to decide when the caller has actually stopped talking, and that trades directly against being interrupted or leaving dead air. Set it aggressively and the agent talks over people. Set it conservatively and the agent feels slow. It is also the one setting I tune per moment rather than globally, because a caller reading out a phone number and a caller saying \"yes\" want opposite values.</p>\n\n<p>Nobody has ever complained that a lecture summary arrived four hundred milliseconds late.</p>\n\n<h2>\n\n\nOne fails loudly and one fails silently\n</h2>\n\n<p>This is the difference I underrate least now.</p>\n\n<p>A bad lecture transcript is obvious. The user reads it, sees the mess, re-records or complains. The feedback loop is short and self-correcting.</p>\n\n<p>A bad phone capture is silent. The call sounds fine. The caller is polite. The agent confirms something plausible, the automation writes a record, and the failure surfaces days later when an email bounces or someone does not turn up to an appointment. Nobody watched it happen.</p>\n\n<p>That is why I score voice agents per field rather than per call, and why I store the raw transcript and audio next to the extracted result. When a structured value is wrong, the only way to find out why is to go back to the signal.</p>\n\n<h2>\n\n\nIf you are choosing between them today\n</h2>\n\n<p>A quick heuristic, since founders ask me this often:</p>\n\n<ul>\n<li>\n<strong>If nobody is waiting, buy accuracy.</strong> Larger model, second pass, domain vocabulary, spend the seconds. Cost per hour of audio is your constraint, not latency.</li>\n<li>\n<strong>If someone is on the line, buy predictability.</strong> Consistent, boring, low latency beats a slightly better transcript, because the caller feels the pause and never sees the transcript.</li>\n<li>\n<strong>Store the raw signal.</strong> Transcript, audio, timestamps. Every interesting debugging session I have had in either product started by going back to the recording.</li>\n</ul>\n\n<p>The uncomfortable conclusion after shipping both is that the model was rarely the interesting variable. Capture conditions, chunking, structured extraction and honest evaluation moved both products more than any provider comparison did, and all four are things you own rather than things you buy.</p>\n\n<p>I wrote a longer version of this on my own site, with more of the voice agent stack around it: <a href=\"https://nabeelbaghoor.com/blog/speech-to-text-lectures-vs-phone-calls/\" rel=\"noopener noreferrer\">Speech-to-text for lectures vs phone calls</a>.</p>","content":"The first time I opened a Retell call transcript, I assumed I already knew this problem.\nI had shipped LectureNotes AI before that: a note-taking app that records a lecture and turns it into a summary and a revision outline. Audio in, text out, users happy. So when voice agents became most of my work, I filed speech recognition under solved and moved on to the parts I thought were hard.\nI was wrong in a way that cost me a few weeks. Transcribing a fifty minute lecture and transcribing a caller on a phone line share a model family and share almost nothing else. They fail differently, they are evaluated differently, and the thing that improves each one is in a completely different layer of the stack.\nHere is the honest split, after shipping both.\nOne is a batch job, one is a turn\nLong-form transcription is a batch problem. A student hits record, sits through an hour, closes the laptop. The audio arrives as one file. Nobody is waiting on a socket for the next token. If a pass comes out badly you can run it again with different settings and nobody knows. You can spend three minutes of compute on an hour of audio and the user experiences that as \"it was ready when I looked\".\nA voice agent inverts every one of those properties. Audio arrives as a stream, the useful unit is a three to eight second turn, and the transcript is consumed immediately by a model that is about to say something out loud to a human being. There is no second pass and no running it again with a bigger model. The caller is on the line, and every millisecond you spend is silence they are listening to.\nOnce you frame it that way, the question stops being which STT provider is best and starts being which failure you can afford.\nIn batch, the capture path beats the model\nThe single biggest accuracy win on the note-taking side had nothing to do with the transcription model. It was recording conditions.\nA lecture recording is a phone lying on a desk six rows back, in a room with hard walls and an air conditioner. The lecturer walks around. Against that, the gap between two good STT models is noise.\nWhat actually helped was unglamorous: handling the recording session properly so it survives the screen locking and the app being backgrounded, being honest in the UI about what a bad recording is going to produce, and getting the sample rate and format right at the source instead of repairing it downstream. If you upsample rubbish you get expensive rubbish.\nIf you are building anything that ingests audio, that is where I would spend the first week, not on provider benchmarks.\nWord error rate measures the wrong words\nWER treats every word as equal. Users do not.\nA lecture transcript that is broadly correct but mangles the lecturer's specific terminology is useless, because the terminology is the entire reason a student is reading it. Course jargon, people's names, formula names, abbreviations: those are exactly the tokens with the least redundancy in the surrounding sentence, so context cannot repair them. Meanwhile a transcript that quietly drops filler words and tidies false starts usually tests better than a faithful one.\nSo the useful evaluation was never an aggregate score. It was: did the terms survive? Feeding known domain vocabulary in as a bias list or prompt hint bought more perceived quality than any model swap I tried, because it targeted the small set of tokens carrying all the meaning.\nThe same shape shows up on the phone side with different tokens. Identifiers have no redundancy either: nothing around a postcode, a registration number or an email address tells the model what it should have been.\nPick the twenty terms or the five fields that actually matter to your user and measure those. An aggregate score will keep telling you things are fine while the only words anyone cares about are wrong.\nWhat happens after the transcript is the actual product\nFor long-form, the transcript is not the deliverable. The summary and the outline are, and that is where the quality is genuinely won.\nThree things mattered more than I expected:\nChunk on the audio, not on the character count. Cutting a transcript into fixed windows regularly slices a definition in half, and the summary then confidently loses it. Chunking around natural pauses, with overlap between chunks, produced noticeably better output for the same model and the same prompt.\nKeep timestamps all the way through. If a student cannot jump from a summary line back to the moment it came from, they cannot check it, and if they cannot check it they will not trust it. Traceability is a feature, not plumbing.\nStructure is a product decision, not a prompt afterthought. Students did not want prose. They wanted something scannable: headings, short lines, definitions pulled out. An excellent essay-shaped summary tested worse than a mediocre outline-shaped one.\nI run the same instinct on voice work now. The transcript is raw signal. Everything the client actually cares about, the booking, the CRM record, the outcome label, is a structured extraction sitting on top of it.\nOn a phone line, you are transcribing through a straw\nThen there is the call.\nTelephony audio is narrowband. This is the part I wish someone had said to me plainly: it is not a lightly degraded version of good audio, it is audio with the top of the spectrum removed, and the frequencies it removes are the ones that distinguish consonants. That is why f and s collapse into each other, why m and n become a coin toss, and why five and nine are a permanent problem when you are taking a phone number. It is not the model failing. That information never reached the model.\nOn top of that you get callers in cars, on speaker, standing in a busy shop, with regional accents that a US-tuned recognition setting handles badly.\nBecause the audio is worse and the stakes are higher, the fix has to leave the STT layer entirely. You cannot make the line better. You can design a conversation that survives mishearing: confirm identifiers in small chunks as you take them, validate shape downstream in automation rather than hoping the prompt caught it, and build a repair ladder with an exit, so a bad capture becomes a text message or a transfer instead of a loop.\nLatency is the product in exactly one of them\nIn batch, latency is a scheduling detail. In a call, it is the product.\nStreaming STT adds cost at the point where you are most sensitive. Endpointing has to decide when the caller has actually stopped talking, and that trades directly against being interrupted or leaving dead air. Set it aggressively and the agent talks over people. Set it conservatively and the agent feels slow. It is also the one setting I tune per moment rather than globally, because a caller reading out a phone number and a caller saying \"yes\" want opposite values.\nNobody has ever complained that a lecture summary arrived four hundred milliseconds late.\nOne fails loudly and one fails silently\nThis is the difference I underrate least now.\nA bad lecture transcript is obvious. The user reads it, sees the mess, re-records or complains. The feedback loop is short and self-correcting.\nA bad phone capture is silent. The call sounds fine. The caller is polite. The agent confirms something plausible, the automation writes a record, and the failure surfaces days later when an email bounces or someone does not turn up to an appointment. Nobody watched it happen.\nThat is why I score voice agents per field rather than per call, and why I store the raw transcript and audio next to the extracted result. When a structured value is wrong, the only way to find out why is to go back to the signal.\nIf you are choosing between them today\nA quick heuristic, since founders ask me this often:\n- If nobody is waiting, buy accuracy. Larger model, second pass, domain vocabulary, spend the seconds. Cost per hour of audio is your constraint, not latency.\n- If someone is on the line, buy predictability. Consistent, boring, low latency beats a slightly better transcript, because the caller feels the pause and never sees the transcript.\n- Store the raw signal. Transcript, audio, timestamps. Every interesting debugging session I have had in either product started by going back to the recording.\nThe uncomfortable conclusion after shipping both is that the model was rarely the interesting variable. Capture conditions, chunking, structured extraction and honest evaluation moved both products more than any provider comparison did, and all four are things you own rather than things you buy.\nI wrote a longer version of this on my own site, with more of the voice agent stack around it: Speech-to-text for lectures vs phone calls.\nTop comments (0)","published":"Sat, 12 Sep 2026 21:44:02 +0000","author":"Nabeel Hassan","guid":"https://dev.to/nabeelbaghoor/transcribing-a-lecture-and-transcribing-a-caller-are-not-the-same-problem-45j0","created_at":"2026-09-13T00:35:43.354162","last_synchronized":"2026-09-13T00:35:43.354162","sentiment":{"sentiment":"Negative","score":-0.6806,"details":{"neg":0.074,"neu":0.853,"pos":0.073,"compound":-0.6806}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"AI Won’t Fix a Broken Business Process","link":"https://dev.to/ikilic/ai-wont-fix-a-broken-business-process-3ma0","description":"<p><em>Why enterprise AI projects should start with workflow design, not model selection</em></p>\n\n<p>An enterprise AI project often starts with the same question:</p>\n\n<p><strong>Which model should we use?</strong></p>\n\n<p>Should it be GPT, Claude, Gemini, an open-source model, or a smaller model running privately?</p>\n\n<p>That question matters, but it is rarely the best starting point.</p>\n\n<p>A more useful question is:</p>\n\n<p><strong>Which business process are we trying to improve?</strong></p>\n\n<p>Many enterprise processes are already difficult before AI enters the picture. Data is incomplete. Responsibilities are unclear. Different systems contain conflicting information. Business rules exist only in people’s memories. Exceptions are handled manually. Nobody can clearly explain where a process begins, where it ends, or how success is measured.</p>\n\n<p>Adding AI to such a process does not automatically solve these problems.</p>\n\n<p>It may simply make the process faster, larger, and more difficult to understand.</p>\n\n<p><strong>1. The AI Project Usually Starts in the Wrong Place</strong><br />\nImagine a company building an AI assistant for its sales team.</p>\n\n<p>The assistant should recommend the next action for each customer. It might suggest a follow-up call, identify an inactive opportunity, or remind a salesperson about an unanswered request.</p>\n\n<p>At first, this sounds like a straightforward AI feature.</p>\n\n<p>But several questions appear immediately:</p>\n\n<p>What does “next action” actually mean? Who is responsible for defining it? Which data should the assistant trust? What happens when the CRM contains duplicate customer records? How should it treat an opportunity with no recent activity? What if the customer has already contacted support about the same issue? How do we measure whether the recommendation was useful?</p>\n\n<p>These are not primarily model-selection questions.</p>\n\n<p>They are questions about process definition, data ownership, business rules, permissions, and workflow design.</p>\n\n<p>A team can spend weeks improving prompts while the real problem remains unresolved. The model may become better at producing recommendations, but the recommendations are still based on unclear inputs and an ambiguous process.</p>\n\n<p>The result is a polished solution to the wrong problem.</p>\n\n<p><strong>2. AI Can Scale a Broken Process</strong><br />\nConsider a typical customer follow-up process.</p>\n\n<p>A salesperson records a meeting in the CRM. The system creates a follow-up task. A manager reviews the opportunity. Another system contains the customer’s payment status. The salesperson is expected to decide what should happen next.</p>\n\n<p>Now imagine adding AI.</p>\n\n<p>The model summarizes the meeting, extracts action items, recommends a follow-up date, and assigns a task automatically.</p>\n\n<p>That may be useful. But it does not resolve the underlying process problems.</p>\n\n<p>If customer ownership is unclear, the AI may assign the task to the wrong person.</p>\n\n<p>If the CRM contains duplicate records, the AI may attach the activity to the wrong customer.</p>\n\n<p>If the business has no clear definition of an inactive opportunity, the model may produce inconsistent recommendations.</p>\n\n<p>If the payment system and CRM disagree, the AI may confidently interpret the wrong status.</p>\n\n<p>The system now operates more quickly, but the underlying confusion remains.</p>\n\n<p>This is more than a traditional “garbage in, garbage out” problem.</p>\n\n<p>A human working with poor information may make one mistake. An automated AI workflow can repeat the same mistake across thousands of records.</p>\n\n<p><strong>A human can create isolated confusion. AI can scale it.</strong></p>\n\n<p>That is why the phrase “chaos in, speed and scale out” describes a real enterprise risk. AI does not need to be inaccurate in every case to cause damage. It only needs to operate inside a poorly defined process without sufficient controls.</p>\n\n<p><strong>3. Before Choosing a Model, Understand the Workflow</strong><br />\nA useful way to examine any business process is to break it into five parts:</p>\n\n<p>Input<br />\n↓<br />\nDecision<br />\n↓<br />\nAction<br />\n↓<br />\nValidation<br />\n↓<br />\nOutcome</p>\n\n<p>Each part deserves separate attention.</p>\n\n<p><strong>Input</strong>: What information enters the process? Where does it come from? Is it complete, current, and correctly associated with the right customer, order, employee, or transaction?</p>\n\n<p><strong>Decision</strong>: What decision must be made? Is it based on explicit business rules, interpretation, or judgment? Can the decision be expressed deterministically?</p>\n\n<p><strong>Action</strong>: What happens after the decision? Is a task created, a message sent, a record updated, or a financial transaction initiated?</p>\n\n<p><strong>Validation</strong>: What must be checked before the action is accepted? Are permissions, limits, approvals, and business constraints enforced?</p>\n\n<p><strong>Outcome</strong>: How do we know the process worked? Did the customer receive the correct response? Was the task completed? Did the action reduce manual work or improve response time?</p>\n\n<p>This decomposition helps identify where AI belongs.</p>\n\n<p>AI may be useful for interpreting an unstructured meeting note, extracting information from an email, classifying a customer request, or suggesting a next action.</p>\n\n<p>It should not automatically replace deterministic validation, authorization, transaction handling, or business rules.</p>\n\n<p>The model can help interpret ambiguity. The application must still control what is allowed to happen.</p>\n\n<p><strong>4. Data Ownership Is Part of the AI Architecture</strong><br />\nData quality is often treated as a preparation step before the AI project begins. In reality, it is part of the architecture itself.</p>\n\n<p>A workflow may depend on a CRM, ERP, ticketing system, internal documents, and communication tools. The important question is not simply whether the data exists.</p>\n\n<p>The important questions are: Which system is authoritative? Who owns each field? How fresh must the data be? How are customers, users, and transactions identified across systems? Which information can the AI access, and under whose permissions?</p>\n\n<p>These are architectural questions because they directly affect what the AI is allowed to interpret and what the application is allowed to do.</p>\n\n<p>A model may generate a convincing recommendation from incomplete or conflicting information. That does not make the recommendation reliable. If two systems disagree about a customer’s status, the AI cannot resolve that conflict merely by producing a more fluent answer. The workflow needs an explicit rule for resolving it.</p>\n\n<p>Data ownership, therefore, is not administrative housekeeping. It is part of the control system that makes AI useful inside an enterprise.</p>\n\n<p><strong>5. Use AI Where It Creates Leverage</strong><br />\nNot every step in a business process needs AI.</p>\n\n<p>Some tasks are already deterministic:</p>\n\n<p>Checking whether a required field is empty<br />\nCalculating a discount<br />\nVerifying a user’s permission<br />\nMatching an exact customer identifier<br />\nApplying a contractual limit<br />\nUpdating a transaction inside a database<br />\nUsing an LLM for these tasks may introduce unnecessary cost, latency, and uncertainty.</p>\n\n<p>Other tasks involve ambiguity:</p>\n\n<p>Understanding the meaning of an email<br />\nExtracting information from a document<br />\nSummarizing a conversation<br />\nClassifying a customer request<br />\nDetecting the likely intent behind a message<br />\nSuggesting a response or next action<br />\nThese are stronger candidates for AI because language models can handle variation and unstructured information more flexibly than traditional rules.</p>\n\n<p>A useful principle is:</p>\n\n<p><strong>Use AI for ambiguity. Use deterministic software for certainty.</strong></p>\n\n<p>The goal is not to introduce the most sophisticated AI architecture into every feature. It is to use an approach that creates meaningful value while keeping failure understandable and controllable.</p>\n\n<p><strong>6. The Real Unit of AI Adoption Is the Workflow</strong><br />\nAn AI feature can look successful in isolation and still fail in practice.</p>\n\n<p>A meeting-summary feature may generate excellent summaries. But what happens afterward?</p>\n\n<p>Does the summary connect to the correct customer? Are action items extracted reliably? Are tasks assigned to the right people? Are deadlines represented correctly? Does the salesperson actually use the generated tasks? Can a manager see whether follow-ups were completed?</p>\n\n<p>The value is not created by the summary alone.</p>\n\n<p>The value is created when the summary becomes part of a functioning workflow.</p>\n\n<p>This changes how AI projects should be measured.</p>\n\n<p>Instead of focusing only on model quality or response fluency, teams should examine operational outcomes:</p>\n\n<p>How much manual work was removed?<br />\nDid response times improve?<br />\nDid routing errors decrease?<br />\nWas less information re-entered across systems?<br />\nDid users accept or ignore the recommendations?<br />\nHow often did humans need to correct the result?<br />\nDid the process create new review or rework costs?<br />\nA model response is an intermediate artifact. The business outcome is the real product.</p>\n\n<p><strong>7. A Practical Way to Start</strong><br />\nA more reliable enterprise AI project can begin with a small process-mapping exercise.</p>\n\n<p>First, document the process as it actually happens — not as it is supposed to happen according to a presentation or specification.</p>\n\n<p>Identify the systems involved, the people responsible, the common exceptions, and the points where manual work or confusion appears.</p>\n\n<p>Then simplify the process before automating it. Remove unnecessary handoffs. Clarify ownership. Define the meaning of important fields. Decide which system is authoritative for each type of information.</p>\n\n<p>Next, separate interpretation from enforcement.</p>\n\n<p>Ask which parts require language understanding or judgment, and which parts should remain deterministic. Define the boundaries before selecting the model.</p>\n\n<p>Only then choose a narrow AI opportunity and define its expected outcome. The first version might only prepare a recommendation for a human rather than execute an action automatically.</p>\n\n<p>For a low-risk feature such as summarization, rewriting, or translation, a team may reasonably test a model quickly. The cost of failure is limited, and the prototype can help determine whether the feature is useful at all.</p>\n\n<p>But once the feature becomes part of a business-critical workflow, process design can no longer be postponed. A successful demo is not evidence that the surrounding business system is ready for automation.</p>\n\n<p>The level of control should reflect the cost of failure.</p>\n\n<p>Low-risk tasks may run with minimal intervention. Higher-impact actions should involve validation, approval, monitoring, or a reliable way to reverse the result.</p>\n\n<p><strong>8. The Model Is Not the Business Process</strong><br />\nThe model is only one component in the system.</p>\n\n<p>The surrounding application still owns the responsibilities that make the workflow dependable: identity, permissions, transactions, business rules, workflow state, retries, idempotency, auditability, error handling, and user experience.</p>\n\n<p>The model may interpret a request, extract information, classify an issue, or suggest the next action. The application must still determine whether that suggestion is permitted, valid, and safe to execute.</p>\n\n<p>This distinction becomes especially important when an AI feature moves from assisting a user to taking action on the user’s behalf.</p>\n\n<p>Sending a draft email and sending a legally significant customer notification are not equivalent operations. Recommending a discount and applying that discount to an order are not equivalent operations either.</p>\n\n<p>The more costly the failure, the stronger the control boundary should be.</p>\n\n<p>That is why the right architecture is rarely “let the model run the process.” It is closer to this:</p>\n\n<p>The model interprets<br />\nThe application validates<br />\nThe business rules decide<br />\nThe workflow executes<br />\nThe system records<br />\nThis does not make the AI less useful. It gives the AI a place where its strengths can be used without allowing its uncertainty to control the entire business process.</p>\n\n<p><strong>Conclusion</strong><br />\nAI does not remove the need for process design.</p>\n\n<p>It makes process design harder to ignore.</p>\n\n<p>When a workflow is unclear, AI may hide the underlying problems behind fluent language and fast execution. When data ownership is weak, AI may spread incorrect interpretations across systems. When business rules are undefined, AI may produce recommendations that sound reasonable but cannot be safely enforced.</p>\n\n<p>The better approach is to start with the workflow.</p>\n\n<p>Understand the inputs. Clarify the decisions. Define the actions. Establish validation. Measure the outcome. Then decide where AI can create real leverage.</p>\n\n<p>The model matters, but it is not the business process.</p>\n\n<p><strong>AI won’t fix a broken business process. But it can make a well-designed process significantly more capable.</strong></p>","content":"Why enterprise AI projects should start with workflow design, not model selection\nAn enterprise AI project often starts with the same question:\nWhich model should we use?\nShould it be GPT, Claude, Gemini, an open-source model, or a smaller model running privately?\nThat question matters, but it is rarely the best starting point.\nA more useful question is:\nWhich business process are we trying to improve?\nMany enterprise processes are already difficult before AI enters the picture. Data is incomplete. Responsibilities are unclear. Different systems contain conflicting information. Business rules exist only in people’s memories. Exceptions are handled manually. Nobody can clearly explain where a process begins, where it ends, or how success is measured.\nAdding AI to such a process does not automatically solve these problems.\nIt may simply make the process faster, larger, and more difficult to understand.\n1. The AI Project Usually Starts in the Wrong Place\nImagine a company building an AI assistant for its sales team.\nThe assistant should recommend the next action for each customer. It might suggest a follow-up call, identify an inactive opportunity, or remind a salesperson about an unanswered request.\nAt first, this sounds like a straightforward AI feature.\nBut several questions appear immediately:\nWhat does “next action” actually mean? Who is responsible for defining it? Which data should the assistant trust? What happens when the CRM contains duplicate customer records? How should it treat an opportunity with no recent activity? What if the customer has already contacted support about the same issue? How do we measure whether the recommendation was useful?\nThese are not primarily model-selection questions.\nThey are questions about process definition, data ownership, business rules, permissions, and workflow design.\nA team can spend weeks improving prompts while the real problem remains unresolved. The model may become better at producing recommendations, but the recommendations are still based on unclear inputs and an ambiguous process.\nThe result is a polished solution to the wrong problem.\n2. AI Can Scale a Broken Process\nConsider a typical customer follow-up process.\nA salesperson records a meeting in the CRM. The system creates a follow-up task. A manager reviews the opportunity. Another system contains the customer’s payment status. The salesperson is expected to decide what should happen next.\nNow imagine adding AI.\nThe model summarizes the meeting, extracts action items, recommends a follow-up date, and assigns a task automatically.\nThat may be useful. But it does not resolve the underlying process problems.\nIf customer ownership is unclear, the AI may assign the task to the wrong person.\nIf the CRM contains duplicate records, the AI may attach the activity to the wrong customer.\nIf the business has no clear definition of an inactive opportunity, the model may produce inconsistent recommendations.\nIf the payment system and CRM disagree, the AI may confidently interpret the wrong status.\nThe system now operates more quickly, but the underlying confusion remains.\nThis is more than a traditional “garbage in, garbage out” problem.\nA human working with poor information may make one mistake. An automated AI workflow can repeat the same mistake across thousands of records.\nA human can create isolated confusion. AI can scale it.\nThat is why the phrase “chaos in, speed and scale out” describes a real enterprise risk. AI does not need to be inaccurate in every case to cause damage. It only needs to operate inside a poorly defined process without sufficient controls.\n3. Before Choosing a Model, Understand the Workflow\nA useful way to examine any business process is to break it into five parts:\nInput\n↓\nDecision\n↓\nAction\n↓\nValidation\n↓\nOutcome\nEach part deserves separate attention.\nInput: What information enters the process? Where does it come from? Is it complete, current, and correctly associated with the right customer, order, employee, or transaction?\nDecision: What decision must be made? Is it based on explicit business rules, interpretation, or judgment? Can the decision be expressed deterministically?\nAction: What happens after the decision? Is a task created, a message sent, a record updated, or a financial transaction initiated?\nValidation: What must be checked before the action is accepted? Are permissions, limits, approvals, and business constraints enforced?\nOutcome: How do we know the process worked? Did the customer receive the correct response? Was the task completed? Did the action reduce manual work or improve response time?\nThis decomposition helps identify where AI belongs.\nAI may be useful for interpreting an unstructured meeting note, extracting information from an email, classifying a customer request, or suggesting a next action.\nIt should not automatically replace deterministic validation, authorization, transaction handling, or business rules.\nThe model can help interpret ambiguity. The application must still control what is allowed to happen.\n4. Data Ownership Is Part of the AI Architecture\nData quality is often treated as a preparation step before the AI project begins. In reality, it is part of the architecture itself.\nA workflow may depend on a CRM, ERP, ticketing system, internal documents, and communication tools. The important question is not simply whether the data exists.\nThe important questions are: Which system is authoritative? Who owns each field? How fresh must the data be? How are customers, users, and transactions identified across systems? Which information can the AI access, and under whose permissions?\nThese are architectural questions because they directly affect what the AI is allowed to interpret and what the application is allowed to do.\nA model may generate a convincing recommendation from incomplete or conflicting information. That does not make the recommendation reliable. If two systems disagree about a customer’s status, the AI cannot resolve that conflict merely by producing a more fluent answer. The workflow needs an explicit rule for resolving it.\nData ownership, therefore, is not administrative housekeeping. It is part of the control system that makes AI useful inside an enterprise.\n5. Use AI Where It Creates Leverage\nNot every step in a business process needs AI.\nSome tasks are already deterministic:\nChecking whether a required field is empty\nCalculating a discount\nVerifying a user’s permission\nMatching an exact customer identifier\nApplying a contractual limit\nUpdating a transaction inside a database\nUsing an LLM for these tasks may introduce unnecessary cost, latency, and uncertainty.\nOther tasks involve ambiguity:\nUnderstanding the meaning of an email\nExtracting information from a document\nSummarizing a conversation\nClassifying a customer request\nDetecting the likely intent behind a message\nSuggesting a response or next action\nThese are stronger candidates for AI because language models can handle variation and unstructured information more flexibly than traditional rules.\nA useful principle is:\nUse AI for ambiguity. Use deterministic software for certainty.\nThe goal is not to introduce the most sophisticated AI architecture into every feature. It is to use an approach that creates meaningful value while keeping failure understandable and controllable.\n6. The Real Unit of AI Adoption Is the Workflow\nAn AI feature can look successful in isolation and still fail in practice.\nA meeting-summary feature may generate excellent summaries. But what happens afterward?\nDoes the summary connect to the correct customer? Are action items extracted reliably? Are tasks assigned to the right people? Are deadlines represented correctly? Does the salesperson actually use the generated tasks? Can a manager see whether follow-ups were completed?\nThe value is not created by the summary alone.\nThe value is created when the summary becomes part of a functioning workflow.\nThis changes how AI projects should be measured.\nInstead of focusing only on model quality or response fluency, teams should examine operational outcomes:\nHow much manual work was removed?\nDid response times improve?\nDid routing errors decrease?\nWas less information re-entered across systems?\nDid users accept or ignore the recommendations?\nHow often did humans need to correct the result?\nDid the process create new review or rework costs?\nA model response is an intermediate artifact. The business outcome is the real product.\n7. A Practical Way to Start\nA more reliable enterprise AI project can begin with a small process-mapping exercise.\nFirst, document the process as it actually happens — not as it is supposed to happen according to a presentation or specification.\nIdentify the systems involved, the people responsible, the common exceptions, and the points where manual work or confusion appears.\nThen simplify the process before automating it. Remove unnecessary handoffs. Clarify ownership. Define the meaning of important fields. Decide which system is authoritative for each type of information.\nNext, separate interpretation from enforcement.\nAsk which parts require language understanding or judgment, and which parts should remain deterministic. Define the boundaries before selecting the model.\nOnly then choose a narrow AI opportunity and define its expected outcome. The first version might only prepare a recommendation for a human rather than execute an action automatically.\nFor a low-risk feature such as summarization, rewriting, or translation, a team may reasonably test a model quickly. The cost of failure is limited, and the prototype can help determine whether the feature is useful at all.\nBut once the feature becomes part of a business-critical workflow, process design can no longer be postponed. A successful demo is not evidence that the surrounding business system is ready for automation.\nThe level of control should reflect the cost of failure.\nLow-risk tasks may run with minimal intervention. Higher-impact actions should involve validation, approval, monitoring, or a reliable way to reverse the result.\n8. The Model Is Not the Business Process\nThe model is only one component in the system.\nThe surrounding application still owns the responsibilities that make the workflow dependable: identity, permissions, transactions, business rules, workflow state, retries, idempotency, auditability, error handling, and user experience.\nThe model may interpret a request, extract information, classify an issue, or suggest the next action. The application must still determine whether that suggestion is permitted, valid, and safe to execute.\nThis distinction becomes especially important when an AI feature moves from assisting a user to taking action on the user’s behalf.\nSending a draft email and sending a legally significant customer notification are not equivalent operations. Recommending a discount and applying that discount to an order are not equivalent operations either.\nThe more costly the failure, the stronger the control boundary should be.\nThat is why the right architecture is rarely “let the model run the process.” It is closer to this:\nThe model interprets\nThe application validates\nThe business rules decide\nThe workflow executes\nThe system records\nThis does not make the AI less useful. It gives the AI a place where its strengths can be used without allowing its uncertainty to control the entire business process.\nConclusion\nAI does not remove the need for process design.\nIt makes process design harder to ignore.\nWhen a workflow is unclear, AI may hide the underlying problems behind fluent language and fast execution. When data ownership is weak, AI may spread incorrect interpretations across systems. When business rules are undefined, AI may produce recommendations that sound reasonable but cannot be safely enforced.\nThe better approach is to start with the workflow.\nUnderstand the inputs. Clarify the decisions. Define the actions. Establish validation. Measure the outcome. Then decide where AI can create real leverage.\nThe model matters, but it is not the business process.\nAI won’t fix a broken business process. But it can make a well-designed process significantly more capable.\nTop comments (0)","published":"Sat, 12 Sep 2026 21:41:33 +0000","author":"ibrahim Kılıç","guid":"https://dev.to/ikilic/ai-wont-fix-a-broken-business-process-3ma0","created_at":"2026-09-13T00:35:43.354162","last_synchronized":"2026-09-13T00:35:43.354162","sentiment":{"sentiment":"Positive","score":0.9966,"details":{"neg":0.076,"neu":0.818,"pos":0.106,"compound":0.9966}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"Why Generic Outreach Fails: Building Custom State Machines to Solve Real Developer Bottlenecks","link":"https://dev.to/nwfd_engineer/why-generic-outreach-fails-building-custom-state-machines-to-solve-real-developer-bottlenecks-40hc","description":"<p>​We’ve all seen the boilerplate comments flooding technical boards—generic copy-pasted text that misses the mark entirely. Scaling effective outreach isn't about spamming; it's about context, precision, and building the right guardrails into your workflow.<br />\n​Over the past few weeks while building custom automation tools, I ran into a subtle but frustrating operational edge case: accidental clipboard cross-posting. When processing leads rapidly, it is entirely too easy to copy a payload meant for one specific technical problem and accidentally paste it into an unrelated thread.<br />\n​To fix this properly, you can't just rely on manual discipline. You need architectural enforcement.<br />\n​The Solution: Scoped State Locks<br />\n​Instead of treating the clipboard globally, we implemented a server-side single-use lock bound exclusively to the dashboard's outreach queue.<br />\n​ID Synchronization: Every generated payload injects an explicit target identifier ([ID: XXX]) into both the UI preview and the response string.<br />\n​Database-Backed State: When a reply is generated and copied, an API route (/api/lead//mark-copied) updates a dedicated tracking column (copied_at) in the database.<br />\n​One-Time Consumption: The UI automatically locks the action button after first use, preventing the exact same context from being accidentally dispatched twice.<br />\n​Isolated Workflows: Standard system-wide clipboard functions remain untouched, ensuring your regular development copy-pasting is never disrupted.<br />\n​Building software that solves actual friction points means designing systems that protect you from human error. If you're building out custom dashboards or looking to streamline high-intent outreach workflows, let's connect.</p>","content":"We’ve all seen the boilerplate comments flooding technical boards—generic copy-pasted text that misses the mark entirely. Scaling effective outreach isn't about spamming; it's about context, precision, and building the right guardrails into your workflow.\nOver the past few weeks while building custom automation tools, I ran into a subtle but frustrating operational edge case: accidental clipboard cross-posting. When processing leads rapidly, it is entirely too easy to copy a payload meant for one specific technical problem and accidentally paste it into an unrelated thread.\nTo fix this properly, you can't just rely on manual discipline. You need architectural enforcement.\nThe Solution: Scoped State Locks\nInstead of treating the clipboard globally, we implemented a server-side single-use lock bound exclusively to the dashboard's outreach queue.\nID Synchronization: Every generated payload injects an explicit target identifier ([ID: XXX]) into both the UI preview and the response string.\nDatabase-Backed State: When a reply is generated and copied, an API route (/api/lead//mark-copied) updates a dedicated tracking column (copied_at) in the database.\nOne-Time Consumption: The UI automatically locks the action button after first use, preventing the exact same context from being accidentally dispatched twice.\nIsolated Workflows: Standard system-wide clipboard functions remain untouched, ensuring your regular development copy-pasting is never disrupted.\nBuilding software that solves actual friction points means designing systems that protect you from human error. If you're building out custom dashboards or looking to streamline high-intent outreach workflows, let's connect.\nFor further actions, you may consider blocking this person and/or reporting abuse\nTop comments (0)","published":"Sat, 12 Sep 2026 21:39:22 +0000","author":"NWFD","guid":"https://dev.to/nwfd_engineer/why-generic-outreach-fails-building-custom-state-machines-to-solve-real-developer-bottlenecks-40hc","created_at":"2026-09-13T00:35:43.354162","last_synchronized":"2026-09-13T00:35:43.354162","sentiment":{"sentiment":"Positive","score":0.8646,"details":{"neg":0.077,"neu":0.812,"pos":0.111,"compound":0.8646}}},{"feed_name":"DEV Community","feed_url":"https://dev.to/feed","title":"Running a nested Proxmox homelab and Docker development on the same Windows machine","link":"https://dev.to/yahavtz/running-a-nested-proxmox-homelab-and-docker-development-on-the-same-windows-machine-44c8","description":"<p><em>Tags: docker, proxmox, homelab, windows</em></p>\n\n<p>I sat down to start a new project and Docker Desktop wouldn't run. I knew why immediately, because I'd already paid this bill once — in the opposite direction.</p>\n\n<p>When I first set up Proxmox nested in VMware, VMware wouldn't run either. I fixed it by turning the Windows hypervisor off so the nested VM could reach VT-x/EPT directly. It worked. What I didn't think about at the time was the other half: Docker Desktop needs that same hypervisor. So I had a working homelab and no development environment, and I'd done it to myself months earlier without noticing.</p>\n\n<p>The two are mutually exclusive on one machine:</p>\n\n<ul>\n<li>\n<strong>Hypervisor on:</strong> Docker Desktop works, nested Proxmox fails.</li>\n<li>\n<strong>Hypervisor off:</strong> Proxmox works, Docker Desktop has no engine.</li>\n</ul>\n\n<p>The switch itself is one line in an elevated prompt, and it needs a reboot either way:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight powershell\"><code><span class=\"c\"># nested Proxmox works, Docker Desktop and WSL2 do not</span><span class=\"w\">\n</span><span class=\"n\">bcdedit</span><span class=\"w\"> </span><span class=\"nx\">/set</span><span class=\"w\"> </span><span class=\"nx\">hypervisorlaunchtype</span><span class=\"w\"> </span><span class=\"nx\">off</span><span class=\"w\">\n\n</span><span class=\"c\"># back to the other side</span><span class=\"w\">\n</span><span class=\"n\">bcdedit</span><span class=\"w\"> </span><span class=\"nx\">/set</span><span class=\"w\"> </span><span class=\"nx\">hypervisorlaunchtype</span><span class=\"w\"> </span><span class=\"nx\">auto</span><span class=\"w\">\n</span></code></pre>\n\n</div>\n\n\n\n<p>I didn't just want them to stop fighting, either. I wanted the new development setup to be able to reach PVE, not merely coexist with it.</p>\n\n<h2>\n\n\nWhat didn't work\n</h2>\n\n<p>I went to ChatGPT first. The answers weren't good enough — they were variations on \"pick one\", and I wanted both.</p>\n\n<p>The idea that actually solved it was mine: build a separate VM whose only job is to be the Docker engine for development. Then PVE keeps direct hardware access, development gets a real Docker daemon, and neither one is Windows' problem any more.</p>\n\n<h2>\n\n\nThe setup\n</h2>\n\n<p><strong>DEV01</strong> — Ubuntu 24.04.5 in VMware, 4 vCPU, ~8 GB RAM, 100 GB disk. Docker Engine 29.8.0, Compose v5.5.1, containerd 2.3.5.</p>\n\n<p>The Windows Docker CLI talks to it through a context pointing at <code>ssh://dev</code>. Docker Desktop stays installed for the CLI only, and never runs.</p>\n\n<p>That's the whole trick. The daemon isn't on Windows, so it doesn't care what Windows did to its hypervisor.</p>\n\n<h3>\n\n\nPasswordless SSH first\n</h3>\n\n<p>The Docker context runs every command over SSH, so key auth has to work before anything else does. From Windows:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight powershell\"><code><span class=\"c\"># create a key if you don't have one</span><span class=\"w\">\n</span><span class=\"n\">ssh-keygen</span><span class=\"w\"> </span><span class=\"nt\">-t</span><span class=\"w\"> </span><span class=\"nx\">ed25519</span><span class=\"w\">\n\n</span><span class=\"c\"># copy the public key to the VM (this is the last time it asks for a password)</span><span class=\"w\">\n</span><span class=\"n\">scp</span><span class=\"w\"> </span><span class=\"s2\">\"</span><span class=\"nv\">$</span><span class=\"nn\">env</span><span class=\"p\">:</span><span class=\"nv\">USERPROFILE</span><span class=\"s2\">\\.ssh\\id_ed25519.pub\"</span><span class=\"w\"> </span><span class=\"err\">&lt;</span><span class=\"nx\">user</span><span class=\"err\">&gt;@&lt;</span><span class=\"nx\">dev01-ip</span><span class=\"err\">&gt;</span><span class=\"p\">:</span><span class=\"nx\">/tmp/master.pub</span><span class=\"w\">\n</span></code></pre>\n\n</div>\n\n\n\n<p>Then on the VM:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight shell\"><code><span class=\"nb\">mkdir</span> <span class=\"nt\">-p</span> ~/.ssh <span class=\"o\">&amp;&amp;</span> <span class=\"nb\">chmod </span>700 ~/.ssh\n<span class=\"nb\">touch</span> ~/.ssh/authorized_keys\n<span class=\"nb\">grep</span> <span class=\"nt\">-qxFf</span> /tmp/master.pub ~/.ssh/authorized_keys <span class=\"o\">||</span> <span class=\"nb\">cat</span> /tmp/master.pub <span class=\"o\">&gt;&gt;</span> ~/.ssh/authorized_keys\n<span class=\"nb\">chmod </span>600 ~/.ssh/authorized_keys\n<span class=\"nb\">rm</span> /tmp/master.pub\n</code></pre>\n\n</div>\n\n\n\n<h3>\n\n\nThen an alias, so the IP appears exactly once\n</h3>\n\n<p>Add this to <code>%USERPROFILE%\\.ssh\\config</code>:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight ssh\"><code><span class=\"k\">Host</span> dev\n<span class=\"k\">HostName</span> &lt;dev01-ip&gt;\n<span class=\"k\">User</span> &lt;user&gt;\n<span class=\"k\">IdentityFile</span> C:/Users/&lt;you&gt;/.ssh/id_ed25519\n<span class=\"k\">IdentitiesOnly</span> <span class=\"no\">yes</span>\n<span class=\"k\">ServerAliveInterval</span> <span class=\"m\">30</span>\n<span class=\"k\">ServerAliveCountMax</span> <span class=\"m\">3</span>\n</code></pre>\n\n</div>\n\n\n\n<p><code>ssh dev</code> should now drop you straight into a shell with no password and no IP. <code>ServerAliveInterval</code> matters more than it looks: without it, a long build over a quiet connection can drop halfway.</p>\n\n<h3>\n\n\nAnd finally the context\n</h3>\n\n\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight powershell\"><code><span class=\"n\">docker</span><span class=\"w\"> </span><span class=\"nx\">context</span><span class=\"w\"> </span><span class=\"nx\">create</span><span class=\"w\"> </span><span class=\"nx\">dev</span><span class=\"w\"> </span><span class=\"nt\">--docker</span><span class=\"w\"> </span><span class=\"s2\">\"host=ssh://dev\"</span><span class=\"w\"> </span><span class=\"nt\">--description</span><span class=\"w\"> </span><span class=\"s2\">\"DEV01 Docker Engine\"</span><span class=\"w\">\n</span><span class=\"n\">docker</span><span class=\"w\"> </span><span class=\"nx\">context</span><span class=\"w\"> </span><span class=\"nx\">use</span><span class=\"w\"> </span><span class=\"nx\">dev</span><span class=\"w\">\n</span><span class=\"n\">docker</span><span class=\"w\"> </span><span class=\"nx\">info</span><span class=\"w\">\n</span></code></pre>\n\n</div>\n\n\n\n<p>If <code>docker info</code> prints the VM's engine version, you're done. Everything below is what that costs you.</p>\n\n<h2>\n\n\nThe side effect nobody mentions\n</h2>\n\n<p>WSL2 stops existing too.</p>\n\n<p>The annoying part is that it doesn't <em>say</em> so. <code>wsl --status</code> still cheerfully reports <code>Default Distribution: Ubuntu, Default Version: 2</code>, so everything looks registered and fine — but nothing can actually run. Anything you were hosting in WSL silently loses its home.</p>\n\n<p>For the real state, ask Windows whether a hypervisor is present at all:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight powershell\"><code><span class=\"p\">(</span><span class=\"n\">Get-CimInstance</span><span class=\"w\"> </span><span class=\"nx\">Win32_ComputerSystem</span><span class=\"p\">)</span><span class=\"o\">.</span><span class=\"nf\">HypervisorPresent</span><span class=\"w\">\n</span></code></pre>\n\n</div>\n\n\n\n<h2>\n\n\nWhat breaks when the daemon isn't local\n</h2>\n\n<p>A remote daemon behaves like a local one until it doesn't. Three things bit me.</p>\n\n<h3>\n\n\n1. Bind mounts resolve on the daemon host\n</h3>\n\n<p>Not where you typed the command. <code>./data:/data</code> quietly means \"a folder on the VM\".</p>\n\n<p>Worse: <code>-v /tmp/script.py:/x</code> made Docker <strong>create a directory</strong> named <code>/tmp/script.py</code> on the VM and mount that, so the container failed with <code>can't find __main__ module</code>. Nothing warns you. The path you meant is on Windows; the path Docker used is on the VM, and Docker invented it for you.</p>\n\n<p>Piping the file over stdin avoids the whole class of problem, because nothing has to exist on the daemon host at all:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight powershell\"><code><span class=\"c\"># instead of -v script.py:/x</span><span class=\"w\">\n</span><span class=\"n\">Get-Content</span><span class=\"w\"> </span><span class=\"nx\">script.py</span><span class=\"w\"> </span><span class=\"o\">|</span><span class=\"w\"> </span><span class=\"n\">docker</span><span class=\"w\"> </span><span class=\"nx\">run</span><span class=\"w\"> </span><span class=\"nt\">-i</span><span class=\"w\"> </span><span class=\"nt\">--rm</span><span class=\"w\"> </span><span class=\"nx\">python:3.11-slim</span><span class=\"w\"> </span><span class=\"nx\">python</span><span class=\"w\"> </span><span class=\"o\">-</span><span class=\"w\">\n\n</span><span class=\"c\"># same idea for a shell script</span><span class=\"w\">\n</span><span class=\"n\">Get-Content</span><span class=\"w\"> </span><span class=\"nx\">setup.sh</span><span class=\"w\"> </span><span class=\"o\">|</span><span class=\"w\"> </span><span class=\"n\">docker</span><span class=\"w\"> </span><span class=\"nx\">run</span><span class=\"w\"> </span><span class=\"nt\">-i</span><span class=\"w\"> </span><span class=\"nt\">--rm</span><span class=\"w\"> </span><span class=\"nx\">ubuntu:24.04</span><span class=\"w\"> </span><span class=\"nx\">bash</span><span class=\"w\"> </span><span class=\"nt\">-s</span><span class=\"w\">\n</span></code></pre>\n\n</div>\n\n\n\n<h3>\n\n\n2. <code>127.0.0.1</code> in a port mapping is the VM's loopback now\n</h3>\n\n<p>Published ports are invisible from Windows until you forward them. Everything starts, the health check inside the VM passes, and your browser on the workstation sees nothing.</p>\n\n<p>An SSH tunnel is the smallest fix, and it keeps both ends on loopback so nothing is exposed to the LAN:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight powershell\"><code><span class=\"n\">ssh</span><span class=\"w\"> </span><span class=\"nt\">-N</span><span class=\"w\"> </span><span class=\"nt\">-o</span><span class=\"w\"> </span><span class=\"nx\">ExitOnForwardFailure</span><span class=\"o\">=</span><span class=\"n\">yes</span><span class=\"w\"> </span><span class=\"se\">`\n</span><span class=\"w\"></span><span class=\"nt\">-L</span><span class=\"w\"> </span><span class=\"nx\">127.0.0.1:5236:127.0.0.1:5236</span><span class=\"w\"> </span><span class=\"se\">`\n</span><span class=\"w\"></span><span class=\"nt\">-L</span><span class=\"w\"> </span><span class=\"nx\">127.0.0.1:8090:127.0.0.1:8090</span><span class=\"w\"> </span><span class=\"nx\">dev</span><span class=\"w\">\n</span></code></pre>\n\n</div>\n\n\n\n<p><code>ExitOnForwardFailure=yes</code> is the part worth copying. Without it, a port that's already taken fails silently and you get a tunnel that's half open — which looks exactly like the app being broken.</p>\n\n<p>I wrapped that in the same script that starts the stack, so the tunnel opens and closes with it.</p>\n\n<h3>\n\n\n3. The build context goes over SSH on every build\n</h3>\n\n<p>Without a <code>.dockerignore</code>, the local virtualenv is uploaded each time you build. With one, my context transfer was <strong>51.64 kB</strong>.</p>\n\n<p>This is the one item on the list that's a straight upgrade once you fix it: the feedback is immediate and visible in the build output.</p>\n\n<h2>\n\n\nHGFS: why the share looked empty\n</h2>\n\n<p>I share the Windows folder into the VM with a VMware shared folder and bind it into the containers.</p>\n\n<p>First attempt: the share mounted fine as my login user, and was completely invisible to the containers. The bind failed, or the path just looked empty.</p>\n\n<p>The reason isn't a VMware quirk — it's FUSE's default. <strong>A FUSE filesystem mounted by an ordinary user is closed to every other uid, including root.</strong> dockerd is root. Containers run as root. So they see nothing.</p>\n\n<p><code>allow_other</code> lifts it, with a catch: a non-root user can only pass <code>allow_other</code> if <code>user_allow_other</code> is uncommented in <code>/etc/fuse.conf</code>, and it's commented out by default on Ubuntu 24.04. Mounting as root from <code>/etc/fstab</code> sidesteps that and survives a reboot:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight conf\"><code>.<span class=\"n\">host</span>:/&lt;<span class=\"n\">share</span>-<span class=\"n\">name</span>&gt; /<span class=\"n\">mnt</span>/&lt;<span class=\"n\">share</span>-<span class=\"n\">name</span>&gt; <span class=\"n\">fuse</span>.<span class=\"n\">vmhgfs</span>-<span class=\"n\">fuse</span> <span class=\"n\">defaults</span>,<span class=\"n\">allow_other</span>,<span class=\"n\">uid</span>=<span class=\"m\">1000</span>,<span class=\"n\">gid</span>=<span class=\"m\">1000</span>,<span class=\"n\">nofail</span> <span class=\"m\">0</span> <span class=\"m\">0</span>\n</code></pre>\n\n</div>\n\n\n\n<p><code>nofail</code> so a missing share never blocks boot.</p>\n\n<p>I recognised the symptom fast, but \"looks empty\" is a miserable error message to debug cold.</p>\n\n<h3>\n\n\nOne guard worth having\n</h3>\n\n<p>When the share isn't mounted, the mount point is just an empty directory. Containers happily write into the VM's local disk, everything looks healthy, and the files go nowhere.</p>\n\n<p>So don't check that the directory exists — check that a file you know is in the share is visible through the mount, before starting anything:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight shell\"><code>ssh dev <span class=\"s1\">'test -f /mnt/&lt;share-name&gt;/README.md'</span> <span class=\"o\">||</span> <span class=\"o\">{</span> <span class=\"nb\">echo</span> <span class=\"s2\">\"share is not mounted\"</span><span class=\"p\">;</span> <span class=\"nb\">exit </span>1<span class=\"p\">;</span> <span class=\"o\">}</span>\n</code></pre>\n\n</div>\n\n\n\n<p>An empty directory and a working share are indistinguishable at a glance, and the failure mode is silent data loss.</p>\n\n<h2>\n\n\nSix things I proved before trusting it\n</h2>\n\n<p>All of these ran as root inside a container with the share bound in, then were checked from Windows.</p>\n\n<div class=\"table-wrapper-paragraph\"><table>\n<thead>\n<tr>\n<th>#</th>\n<th>Check</th>\n<th>Result</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>Write a file, read it from Windows</td>\n<td>Works, owner is the Windows user</td>\n</tr>\n<tr>\n<td>2</td>\n<td>Replace a file by rename</td>\n<td>Works, both in Python and .NET, no temp files left</td>\n</tr>\n<tr>\n<td>3</td>\n<td>Set a file's mtime</td>\n<td>Works, exact to the tick</td>\n</tr>\n<tr>\n<td>4</td>\n<td>Delete a file</td>\n<td>Works</td>\n</tr>\n<tr>\n<td>5</td>\n<td>Write a 10 MB image</td>\n<td>Works, write + fsync in 0.014 s, SHA-256 identical both sides</td>\n</tr>\n<tr>\n<td>6</td>\n<td>See a Windows-created file from the container</td>\n<td>No measurable delay</td>\n</tr>\n</tbody>\n</table></div>\n\n<p>Measuring #6 honestly needs the clock offset first, and the way to get it is the same trick NTP uses: take a timestamp on Windows, ask the VM for its time, take a second timestamp on Windows, and compare the VM's answer against the midpoint of the two.<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight powershell\"><code><span class=\"nv\">$t1</span><span class=\"w\"> </span><span class=\"o\">=</span><span class=\"w\"> </span><span class=\"p\">[</span><span class=\"n\">DateTimeOffset</span><span class=\"p\">]::</span><span class=\"n\">UtcNow</span><span class=\"w\">\n</span><span class=\"nv\">$dev</span><span class=\"w\"> </span><span class=\"o\">=</span><span class=\"w\"> </span><span class=\"p\">[</span><span class=\"n\">DateTimeOffset</span><span class=\"p\">]::</span><span class=\"n\">Parse</span><span class=\"p\">((</span><span class=\"n\">ssh</span><span class=\"w\"> </span><span class=\"nx\">dev</span><span class=\"w\"> </span><span class=\"s1\">'date -u --iso-8601=ns'</span><span class=\"p\">))</span><span class=\"w\">\n</span><span class=\"nv\">$t3</span><span class=\"w\"> </span><span class=\"o\">=</span><span class=\"w\"> </span><span class=\"p\">[</span><span class=\"n\">DateTimeOffset</span><span class=\"p\">]::</span><span class=\"n\">UtcNow</span><span class=\"w\">\n</span><span class=\"p\">(</span><span class=\"nv\">$dev</span><span class=\"w\"> </span><span class=\"o\">-</span><span class=\"w\"> </span><span class=\"p\">[</span><span class=\"n\">DateTimeOffset</span><span class=\"p\">]::</span><span class=\"n\">FromUnixTimeMilliseconds</span><span class=\"p\">(</span><span class=\"w\">\n</span><span class=\"p\">(</span><span class=\"nv\">$t1</span><span class=\"o\">.</span><span class=\"nf\">ToUnixTimeMilliseconds</span><span class=\"p\">()</span><span class=\"w\"> </span><span class=\"o\">+</span><span class=\"w\"> </span><span class=\"nv\">$t3</span><span class=\"o\">.</span><span class=\"nf\">ToUnixTimeMilliseconds</span><span class=\"p\">())</span><span class=\"w\"> </span><span class=\"nx\">/</span><span class=\"w\"> </span><span class=\"nx\">2</span><span class=\"p\">))</span><span class=\"o\">.</span><span class=\"nf\">TotalSeconds</span><span class=\"w\">\n</span></code></pre>\n\n</div>\n\n\n\n<p>Run it three times and take the spread as your error bar. Mine came out at <strong>+0.39 s ± 0.1 s</strong> — the VM runs ahead of the workstation — so any latency claim smaller than that is noise, not a measurement. It's worth doing this before you quote yourself a number you'll later repeat.</p>\n\n<h3>\n\n\nThe one real limitation\n</h3>\n\n<p>Appending to a file that <strong>already existed</strong> took ≈1.4 s to become visible in the container — the FUSE attribute cache. Creating, renaming and deleting are immediate.</p>\n\n<p>So: fine if every writer creates a new file or renames one into place. A problem if anything watches files, or re-reads one mid-edit.</p>\n\n<p>Also worth knowing: metadata over HGFS is synthetic. Files always report uid/gid 1000 and mode 0777, and <code>chown</code>/<code>chmod</code> succeed while changing nothing. If your container logic asserts on permissions, it's asserting on a fiction.</p>\n\n<h2>\n\n\nDocker Desktop takes the context back\n</h2>\n\n<p>Mid-session the active context flipped from my SSH context back to <code>desktop-linux</code> on its own. Every command after that failed — Desktop had grabbed the context and then couldn't start, because the hypervisor is off.</p>\n\n<p>The fix is to stop relying on the active context at all:<br />\n</p>\n\n<div class=\"highlight js-code-highlight\">\n<pre class=\"highlight powershell\"><code><span class=\"n\">docker</span><span class=\"w\"> </span><span class=\"nt\">-c</span><span class=\"w\"> </span><span class=\"nx\">dev</span><span class=\"w\"> </span><span class=\"nx\">compose</span><span class=\"w\"> </span><span class=\"nx\">up</span><span class=\"w\"> </span><span class=\"nt\">-d</span><span class=\"w\">\n</span></code></pre>\n\n</div>\n\n\n\n<p>Pass <code>-c &lt;context&gt;</code> in every script, or set <code>DOCKER_CONTEXT</code> for the session. If the context is an implicit global, something else will eventually change it for you.</p>\n\n<h2>\n\n\nIs this permanent?\n</h2>\n\n<p>No. DEV01 gets deleted when PVE moves to dedicated hardware — but that's months away.</p>\n\n<p>PVE still has to prove itself first. The plan is to buy smart-home hardware, let PVE run that, and see whether it earns a server of its own. Until then, this is a good enough answer for months of development, and \"good enough for months\" is a real category.</p>\n\n<h2>\n\n\nWhat I'd tell someone hitting this\n</h2>\n\n<p>Be creative. If Windows can't solve a problem, an extra VM is a legitimate answer, not a defeat.</p>\n\n<p>The framing I was handed everywhere — including by the chatbot — was <em>choose one</em>. The constraint was real; the conclusion wasn't. Moving the daemon out of the argument entirely cost me one VM, one <code>.dockerignore</code>, one fstab line, and an afternoon of proving the shared folder does what I think it does.</p>\n\n\n\n\n<p>I'm Yahav Tzukerman, a full-stack developer (Angular + .NET). I build things I need and write about what broke along the way — most of it lives in my homelab.</p>\n\n<p>I also build automation for small businesses: Telegram bots, document workflows, and AI agents that handle the repetitive parts. If you've got a process that's eating your week, I'm happy to talk about it.</p>\n\n<p>Find me on <a href=\"https://dev.to/yahavtz\">Dev.to</a> — or drop a comment below, I answer all of them.</p>","content":"Tags: docker, proxmox, homelab, windows\nI sat down to start a new project and Docker Desktop wouldn't run. I knew why immediately, because I'd already paid this bill once — in the opposite direction.\nWhen I first set up Proxmox nested in VMware, VMware wouldn't run either. I fixed it by turning the Windows hypervisor off so the nested VM could reach VT-x/EPT directly. It worked. What I didn't think about at the time was the other half: Docker Desktop needs that same hypervisor. So I had a working homelab and no development environment, and I'd done it to myself months earlier without noticing.\nThe two are mutually exclusive on one machine:\n- Hypervisor on: Docker Desktop works, nested Proxmox fails.\n- Hypervisor off: Proxmox works, Docker Desktop has no engine.\nThe switch itself is one line in an elevated prompt, and it needs a reboot either way:\n# nested Proxmox works, Docker Desktop and WSL2 do not\nbcdedit /set hypervisorlaunchtype off\n# back to the other side\nbcdedit /set hypervisorlaunchtype auto\nI didn't just want them to stop fighting, either. I wanted the new development setup to be able to reach PVE, not merely coexist with it.\nWhat didn't work\nI went to ChatGPT first. The answers weren't good enough — they were variations on \"pick one\", and I wanted both.\nThe idea that actually solved it was mine: build a separate VM whose only job is to be the Docker engine for development. Then PVE keeps direct hardware access, development gets a real Docker daemon, and neither one is Windows' problem any more.\nThe setup\nDEV01 — Ubuntu 24.04.5 in VMware, 4 vCPU, ~8 GB RAM, 100 GB disk. Docker Engine 29.8.0, Compose v5.5.1, containerd 2.3.5.\nThe Windows Docker CLI talks to it through a context pointing at ssh://dev\n. Docker Desktop stays installed for the CLI only, and never runs.\nThat's the whole trick. The daemon isn't on Windows, so it doesn't care what Windows did to its hypervisor.\nPasswordless SSH first\nThe Docker context runs every command over SSH, so key auth has to work before anything else does. From Windows:\n# create a key if you don't have one\nssh-keygen -t ed25519\n# copy the public key to the VM (this is the last time it asks for a password)\nscp \"$env:USERPROFILE\\.ssh\\id_ed25519.pub\" <user>@<dev01-ip>:/tmp/master.pub\nThen on the VM:\nmkdir -p ~/.ssh && chmod 700 ~/.ssh\ntouch ~/.ssh/authorized_keys\ngrep -qxFf /tmp/master.pub ~/.ssh/authorized_keys || cat /tmp/master.pub >> ~/.ssh/authorized_keys\nchmod 600 ~/.ssh/authorized_keys\nrm /tmp/master.pub\nThen an alias, so the IP appears exactly once\nAdd this to %USERPROFILE%\\.ssh\\config\n:\nHost dev\nHostName <dev01-ip>\nUser <user>\nIdentityFile C:/Users/<you>/.ssh/id_ed25519\nIdentitiesOnly yes\nServerAliveInterval 30\nServerAliveCountMax 3\nssh dev\nshould now drop you straight into a shell with no password and no IP. ServerAliveInterval\nmatters more than it looks: without it, a long build over a quiet connection can drop halfway.\nAnd finally the context\ndocker context create dev --docker \"host=ssh://dev\" --description \"DEV01 Docker Engine\"\ndocker context use dev\ndocker info\nIf docker info\nprints the VM's engine version, you're done. Everything below is what that costs you.\nThe side effect nobody mentions\nWSL2 stops existing too.\nThe annoying part is that it doesn't say so. wsl --status\nstill cheerfully reports Default Distribution: Ubuntu, Default Version: 2\n, so everything looks registered and fine — but nothing can actually run. Anything you were hosting in WSL silently loses its home.\nFor the real state, ask Windows whether a hypervisor is present at all:\n(Get-CimInstance Win32_ComputerSystem).HypervisorPresent\nWhat breaks when the daemon isn't local\nA remote daemon behaves like a local one until it doesn't. Three things bit me.\n1. Bind mounts resolve on the daemon host\nNot where you typed the command. ./data:/data\nquietly means \"a folder on the VM\".\nWorse: -v /tmp/script.py:/x\nmade Docker create a directory named /tmp/script.py\non the VM and mount that, so the container failed with can't find __main__ module\n. Nothing warns you. The path you meant is on Windows; the path Docker used is on the VM, and Docker invented it for you.\nPiping the file over stdin avoids the whole class of problem, because nothing has to exist on the daemon host at all:\n# instead of -v script.py:/x\nGet-Content script.py | docker run -i --rm python:3.11-slim python -\n# same idea for a shell script\nGet-Content setup.sh | docker run -i --rm ubuntu:24.04 bash -s\n2. 127.0.0.1\nin a port mapping is the VM's loopback now\nPublished ports are invisible from Windows until you forward them. Everything starts, the health check inside the VM passes, and your browser on the workstation sees nothing.\nAn SSH tunnel is the smallest fix, and it keeps both ends on loopback so nothing is exposed to the LAN:\nssh -N -o ExitOnForwardFailure=yes `\n-L 127.0.0.1:5236:127.0.0.1:5236 `\n-L 127.0.0.1:8090:127.0.0.1:8090 dev\nExitOnForwardFailure=yes\nis the part worth copying. Without it, a port that's already taken fails silently and you get a tunnel that's half open — which looks exactly like the app being broken.\nI wrapped that in the same script that starts the stack, so the tunnel opens and closes with it.\n3. The build context goes over SSH on every build\nWithout a .dockerignore\n, the local virtualenv is uploaded each time you build. With one, my context transfer was 51.64 kB.\nThis is the one item on the list that's a straight upgrade once you fix it: the feedback is immediate and visible in the build output.\nHGFS: why the share looked empty\nI share the Windows folder into the VM with a VMware shared folder and bind it into the containers.\nFirst attempt: the share mounted fine as my login user, and was completely invisible to the containers. The bind failed, or the path just looked empty.\nThe reason isn't a VMware quirk — it's FUSE's default. A FUSE filesystem mounted by an ordinary user is closed to every other uid, including root. dockerd is root. Containers run as root. So they see nothing.\nallow_other\nlifts it, with a catch: a non-root user can only pass allow_other\nif user_allow_other\nis uncommented in /etc/fuse.conf\n, and it's commented out by default on Ubuntu 24.04. Mounting as root from /etc/fstab\nsidesteps that and survives a reboot:\n.host:/<share-name> /mnt/<share-name> fuse.vmhgfs-fuse defaults,allow_other,uid=1000,gid=1000,nofail 0 0\nnofail\nso a missing share never blocks boot.\nI recognised the symptom fast, but \"looks empty\" is a miserable error message to debug cold.\nOne guard worth having\nWhen the share isn't mounted, the mount point is just an empty directory. Containers happily write into the VM's local disk, everything looks healthy, and the files go nowhere.\nSo don't check that the directory exists — check that a file you know is in the share is visible through the mount, before starting anything:\nssh dev 'test -f /mnt/<share-name>/README.md' || { echo \"share is not mounted\"; exit 1; }\nAn empty directory and a working share are indistinguishable at a glance, and the failure mode is silent data loss.\nSix things I proved before trusting it\nAll of these ran as root inside a container with the share bound in, then were checked from Windows.\n| # | Check | Result |\n|---|---|---|\n| 1 | Write a file, read it from Windows | Works, owner is the Windows user |\n| 2 | Replace a file by rename | Works, both in Python and .NET, no temp files left |\n| 3 | Set a file's mtime | Works, exact to the tick |\n| 4 | Delete a file | Works |\n| 5 | Write a 10 MB image | Works, write + fsync in 0.014 s, SHA-256 identical both sides |\n| 6 | See a Windows-created file from the container | No measurable delay |\nMeasuring #6 honestly needs the clock offset first, and the way to get it is the same trick NTP uses: take a timestamp on Windows, ask the VM for its time, take a second timestamp on Windows, and compare the VM's answer against the midpoint of the two.\n$t1 = [DateTimeOffset]::UtcNow\n$dev = [DateTimeOffset]::Parse((ssh dev 'date -u --iso-8601=ns'))\n$t3 = [DateTimeOffset]::UtcNow\n($dev - [DateTimeOffset]::FromUnixTimeMilliseconds(\n($t1.ToUnixTimeMilliseconds() + $t3.ToUnixTimeMilliseconds()) / 2)).TotalSeconds\nRun it three times and take the spread as your error bar. Mine came out at +0.39 s ± 0.1 s — the VM runs ahead of the workstation — so any latency claim smaller than that is noise, not a measurement. It's worth doing this before you quote yourself a number you'll later repeat.\nThe one real limitation\nAppending to a file that already existed took ≈1.4 s to become visible in the container — the FUSE attribute cache. Creating, renaming and deleting are immediate.\nSo: fine if every writer creates a new file or renames one into place. A problem if anything watches files, or re-reads one mid-edit.\nAlso worth knowing: metadata over HGFS is synthetic. Files always report uid/gid 1000 and mode 0777, and chown\n/chmod\nsucceed while changing nothing. If your container logic asserts on permissions, it's asserting on a fiction.\nDocker Desktop takes the context back\nMid-session the active context flipped from my SSH context back to desktop-linux\non its own. Every command after that failed — Desktop had grabbed the context and then couldn't start, because the hypervisor is off.\nThe fix is to stop relying on the active context at all:\ndocker -c dev compose up -d\nPass -c <context>\nin every script, or set DOCKER_CONTEXT\nfor the session. If the context is an implicit global, something else will eventually change it for you.\nIs this permanent?\nNo. DEV01 gets deleted when PVE moves to dedicated hardware — but that's months away.\nPVE still has to prove itself first. The plan is to buy smart-home hardware, let PVE run that, and see whether it earns a server of its own. Until then, this is a good enough answer for months of development, and \"good enough for months\" is a real category.\nWhat I'd tell someone hitting this\nBe creative. If Windows can't solve a problem, an extra VM is a legitimate answer, not a defeat.\nThe framing I was handed everywhere — including by the chatbot — was choose one. The constraint was real; the conclusion wasn't. Moving the daemon out of the argument entirely cost me one VM, one .dockerignore\n, one fstab line, and an afternoon of proving the shared folder does what I think it does.\nI'm Yahav Tzukerman, a full-stack developer (Angular + .NET). I build things I need and write about what broke along the way — most of it lives in my homelab.\nI also build automation for small businesses: Telegram bots, document workflows, and AI agents that handle the repetitive parts. If you've got a process that's eating your week, I'm happy to talk about it.\nFind me on Dev.to — or drop a comment below, I answer all of them.\nTop comments (0)","published":"Sat, 12 Sep 2026 21:35:44 +0000","author":"yahav tzukerman","guid":"https://dev.to/yahavtz/running-a-nested-proxmox-homelab-and-docker-development-on-the-same-windows-machine-44c8","created_at":"2026-09-13T00:35:43.354162","last_synchronized":"2026-09-13T00:35:43.354162","sentiment":{"sentiment":"Positive","score":0.979,"details":{"neg":0.044,"neu":0.901,"pos":0.055,"compound":0.979}}}]}