HomeArticles

Articles on AI, engineering and complex systems

Original writing, published when there is something worth saying. Every analysis separates what was checked against the primary source from what is my own reading — and says which is which.

September 4, 2026

A mesma amostra de skills dá 6,57% ou 48,71%, conforme o scanner

Circulou que 48% das skills de agentes de IA são inseguras. A fonte não é um estudo: é o post de um blogueiro de SEO, cujo filtro descartou 261.451 achados para chegar lá. Dois meses depois, NVIDIA e OpenClaw Foundation mediram 67.453 versões com três scanners ao mesmo tempo — e os três devolveram 6,57%, 7,75% e 48,71%.

#ia#agentes#ciberseguranca#fact-check#metodologia
Deepfake's 7 core statistics all come from the people selling the detector — the only official damage figure counts "mentions of AI", and the science says 55%, a coin flip
August 28, 2026

Deepfake's 7 core statistics all come from the people selling the detector — the only official damage figure counts "mentions of AI", and the science says 55%, a coin flip

The most complete page of deepfake statistics I could find warns that the numbers come from the people who sell detectors — and then draws its seven core figures from exactly there. I checked them one by one: Signicat, Entrust twice, Sumsub twice, Resemble AI and iProov. Seven out of seven vendors, no official body, no academia, no public survey. Four rows of the scoreboard are wrong, and the legal errors all point the same way: they make regulation look further along than it is. So I went looking for what actually exists: the only official damage figure is US$ 893 million — 4.28% of the year's losses, under a label the FBI defines as "contains a reference to artificial intelligence"; the only independent measurement of human detection ability is academic and lands at 55.5%, with a confidence interval that crosses 50%, which is a coin flip; and the best public data in the world on this is Brazilian, from Cetic.br, because it tests instead of asking (41% say they are confident, 17% did well on the test, and the two are uncorrelated). I counted the Resemble AI dashboard myself: of the 2,266 incidents, the famous "US$ 1.3 billion" describes 159.

#deepfake#estatisticas#desinformacao#metodologia#fraude#brasil
The cyber defense letter signed by 155 companies contains not one commitment
August 28, 2026

The cyber defense letter signed by 155 companies contains not one commitment

On 27 August 2026 an open letter convened and hosted by OpenAI called for a collective response to AI-enabled cyber attacks, and the press covered the number: more than a hundred companies. I went after the boring question — what exactly did anyone commit to doing. I pulled the four blocks of asks out of a pinned capture of the page and searched inside them for the five signals that separate a commitment from a statement of intent: an amount, a deadline, a binding verb, a named party, a verifiable target. Eighteen imperative sentences and, across twenty cells, not one hit. The amber you see in the figure is a concession I made by hand against my own argument, and I explain why. To prove the ruler measures, I ran the same test on a page by OpenAI itself, from February, which lights three of the five columns: the company writes an amount when it wants to. I also show that the list of signatories changed four times in forty-seven hours — 116, 127, 128, 155 — while the text of the four blocks did not change a single byte, that one company was removed without explanation, and that the page carries two lists whose counts collide on the same number. And I separate, carefully, what is measured from what is my own analysis.

#ia#ciberseguranca#openai#politica-de-tecnologia#verificacao
The chart that tells you which AI runs on your graphics card gets almost everything right — and gets wrong the one calculation that decides
August 27, 2026

The chart that tells you which AI runs on your graphics card gets almost everything right — and gets wrong the one calculation that decides

A chart in Spanish settles in a single table what runs in each tier of video memory, from 4 GB to 256 GB. I redid the math: of the seventeen hardware verdicts, sixteen are right, and every model it cites really exists — the one piece that does not is the RTX 5080 Super, postponed indefinitely because the 3 GB GDDR7 module costs three times the 2 GB one. The error that matters is a different one, and it is methodological: the chart budgets only the weight file and ignores the context cache, which grows while you talk. With both parts measured — weights from the published GGUF, cache calculated from each model's config.json — one of the ten rungs does not add up even in a 32,000-token working session, and it is precisely the most popular one, the 8 GB rung; six of the ten do not add up at the maximum context of the very model the chart recommends. I also separate out four naming traps, among them a "Q8" that has 4.3 bits per weight and a format that is not from the same ecosystem as the others, and three caveats that apply to the whole table, including why two 32 GB cards are not one 64 GB card.

#ia#llm#hardware#quantizacao#vram#didatico
Only 8 of the 20 most-cited OpenAI numbers are from OpenAI — the rest is leaked, a target, or untraceable
August 26, 2026

Only 8 of the 20 most-cited OpenAI numbers are from OpenAI — the rest is leaked, a target, or untraceable

I classified the 20 most-cited numbers about OpenAI: 8 are official, 6 are reporting, 3 are targets, and 3 no one knows where they came from. The 92% is 31 months old. I went to the source of every number circulating in roundups about the company: the ones from OpenAI itself check out — and they are the least interesting; what makes headlines is what it never signed off on. The "92% of the Fortune 500" came from OpenAI's response to the New York Times lawsuit, in January 2024, and circulates with no date; US$ 24bn in revenue is official, US$ 40bn is an estimate, US$ 280bn is a target — lumping the three together with the same verb is wrong even with every number right; SoftBank invested US$ 64.6bn, not "more than 71," and Stargate's US$ 500bn is a compute commitment, not investment. Altman's "US$ 76,001 salary" is the sum of two Form 990 columns, the following year comes to US$ 113,674, and he is tenth of twelve names in compensation. In Brazil, the "50 million users" figure appears identical in August 2025 and August 2026 — messages per day rose 54%.

#ia#openai#estatisticas#metodologia
How many language models exist? 274, 95, or 403,420 — and the coding score everyone cites comes from a test its own creator abandoned
August 26, 2026

How many language models exist? 274, 95, or 403,420 — and the coding score everyone cites comes from a test its own creator abandoned

274, 95, or 403,420 language models: all three counts are correct. And the coding benchmark everyone cites was abandoned by its own creator in February. I went to check the LLM statistics that circulate in roundups: "how many models exist" has no answer, because there is too much definition — 403,420 repositories on Hugging Face (my own measurement, a one-line command), 95 notable models in Stanford's AI Index, 274 in an editorial catalog. The score a lab gives itself held up under checking (zero to four points of difference where the same model was measured from outside); what does not hold up is the ruler: OpenAI declared SWE-bench Verified contaminated and stopped publishing it on February 23, 2026, and in August it is still the test for "who codes best." Version-to-version comparisons get the sign wrong (August's DeepSeek is 8 points ahead of Opus 4.8, not behind); "the most expensive" costs US$ 30 in the roundups and US$ 150 in the real catalog; the largest models with a known size are all open. In Brazil, Maritaca prices in reais and publishes the only suite-cost comparison in national currency: R$ 206 on Sabiá 4 Thinking against R$ 590 on Opus 4.8.

#ia#llm#estatisticas#benchmarks#metodologia
Eu estava medindo recusas invisíveis. A recusa apareceu — e não era invisível
August 26, 2026

Eu estava medindo recusas invisíveis. A recusa apareceu — e não era invisível

Um estudo meu sobre recusas de IA que passam despercebidas dentro de sistemas de agentes foi interrompido por uma recusa: o classificador de salvaguardas bloqueou a geração do corpus, porque um conjunto de prompts SOBRE recusas lê, para um classificador, como material ofensivo. Havia dois caminhos — reescrever o pedido até passar, o que quase sempre funciona, ou parar. Reformular um pedido porque ele foi sinalizado é evasão de salvaguarda, um andar abaixo do jailbreak e da mesma família; um pesquisador que contorna o classificador para estudar o classificador contaminou o próprio objeto. Congelei o braço do estudo com data no arquivo de estado do projeto e me candidatei ao Cyber Verification Program da Anthropic, o canal formal para trabalho de uso duplo com propósito defensivo. A aprovação saiu dentro do prazo de dois dias úteis. O que ela é: uso duplo deixa de ser bloqueado por padrão, dentro do caso de uso submetido e sob monitoramento contínuo. O que ela não é: parceria, certificação ou endosso — uso proibido continua bloqueado com programa ou sem. E fica a lição que o incidente entrega de graça, que é a tese do estudo: uma recusa só é gerenciável quando é legível. A que me bloqueou tinha texto, categoria e porta de saída; as que eu estou medindo chegam ao orquestrador como resultado vazio e são tratadas como sucesso.

#ia#agentes#ciberseguranca#anthropic#pesquisa#salvaguardas
No, nobody proved that Meta reads your WhatsApp. What I found is worse
August 26, 2026

No, nobody proved that Meta reads your WhatsApp. What I found is worse

A lawyer lost ten years of conversations in a single morning and a video concluded that Meta had read what he wrote. I went to check that accusation against WhatsApp technical whitepapers from 2016 to 2026, the client code, archived captures of the Instagram help centre, the 2017 European Commission decision and the dockets of lawsuits in Brazil, the United States and India. The video is wrong for two reasons, and the second buries the argument: banning en masse and noisily is the signature of an automated classifier, not of someone reading — whoever holds a valuable secret capability protects the capability, not the individual case. But what is left in its place is worse. The company has come to define on its own what counts as a protected conversation and to list exceptions it acknowledges itself; no observation available to the public distinguishes a company that cannot read from one that can and does not say so; and the three choices that produce that impossibility — a closed app, no reproducible build, no external audit — are its own, and reversible by it. On 8 May 2026 Meta switched off end-to-end encryption for Instagram messages: the guarantee we were sold as mathematics was always a corporate promise, and corporate promises are revoked.

#criptografia#whatsapp#instagram#meta#privacidade#verificacao
"95% of AI pilots fail" came from 52 interviews — and the three-question test that catches the next one
August 26, 2026

"95% of AI pilots fail" came from 52 interviews — and the three-question test that catches the next one

The "95% of AI pilots fail" figure came from 52 interviews and never measured agents. I went to the primary sources: every agent statistic is a forecast, a self-report or a measurement — and each fails in its own way. Gartner publishes a forecast and a webinar poll on the same page; the AI Index says 70% in its summary and 79% in its own chart; METR clocked developers 19% slower while they believed they were 20% faster, and in 2026 could not repeat the trial because nobody agrees to work without AI; the consistency metric vanished from leaderboards; a record was withdrawn for answer leakage. I measured the plumbing (MCP SDK: 1,087x in 18 months) and Brazil (17% of companies use AI; 68% of that is workflow automation). None of the numbers is false — all of them change weight once the label travels with them.

#ia#agentes#estatisticas#benchmarks#metodologia
Theory of Constraints: the constraint is a place, not an effort
August 25, 2026

Theory of Constraints: the constraint is a place, not an effort

Nathan Barry summed up the Theory of Constraints in one sentence — effort outside the bottleneck makes the bottleneck worse — resting on the Tiago Forte series that circulates as the explanation of the subject. I went to check the series, the book and the sentence, and wrote a queue simulator to measure instead of arguing by analogy. The series has 11 posts, not 3, and the ones that say what to do after you find the bottleneck sit behind a paywall; Goldratt's five steps are not in The Goal but in a 1990 book almost nobody opens; and the moral sentence gets the sign right and the degree wrong: doubling the capacity of whoever is not the constraint did not take out one more item (-0.3%, noise) and made the waiting grow 33% faster — it worsens the waiting, not the output. Elevating the constraint by 25% returned 24.9%. Goldratt's rope costs 5% of throughput and buys a lead time 8,900 times smaller. I climb the ladder in four steps — the bakery, the names, Little's Law and the five steps read at the source, the bridge to AI agents, where the bottleneck moves around — and close by checking Ford, Spanx and Kit. Ten figures of my own, made in code; the same bakery comes back with the data from each experiment.

#teoria-das-restricoes#goldratt#gestao#filas#lei-de-little#agentes-de-ia#didatico
August 25, 2026

The viral chart crowns the Mac, the table crowns the Spark: I measured 43 thousand calls and the winner is no machine at all

Two viral images compare the Mac Studio M5 Ultra 256 GB, RTX 5090, RTX PRO 6000 and DGX Spark on "value per dollar" for running AI at home. I went to check: the metric multiplies stock by flow, the Spark's price died in February and the PRO 6000 came in without a host PC. Then I redid the math that matters — tokens per dollar against the API for the same model — and its first version assumed a person uses the machine 1% of the time. I measured 43,593 real coding-agent calls on this machine: the 1% describes nobody (chat sits near 0.03%; an agent, between 10% and 30%), 96% of the input is context re-reading, and changing the billing rule for that re-reading moves the bill 16 times, while changing machine moves it 3.6. In the one regime where the cheapest machine wins (1.5x), its peak day does not fit in the day and the mean call, at 151.9 thousand tokens, does not fit in the model's context. Buy for sovereignty, not for savings.

#ia#llm#hardware#custo#inferencia-local
The security score belongs to the pair, not the model — and the lab itself had already corrected it
August 25, 2026

The security score belongs to the pair, not the model — and the lab itself had already corrected it

A code security benchmark measured Claude Fable 5 at 59.8% functional correctness and 19.0% security correctness, and the number became a headline about a model that disappointed. Six days later, the same lab, the same author and the same benchmark published the same Fable 5 at 72.6% and 29.0% — the best security score on the table at that moment. They did not change the model; they changed the tool driving it. I went after both texts and found a third, by the same author on the same day, that nobody cites: the benchmark's own anti-cheating audit, which knocked 9 percentage points of security off one combination without anything changing in the model. I show the narrow conclusion the data supports — the score belongs to the harness-and-model pair and to the version of the ruler on the day it ran — its limit (across the eight pairs on the leaderboard the median difference is 1.65 points, and the Fable 5 case is six times that), the Hacker News counterpoint, which accuses the benchmark of being crooked AGAINST the model, and the mechanical finding: I counted the links between the four pages and the graph is one-way — whoever arrives via the text that circulated has no path to the correction. Six original figures, made in code, and every calculation scripted over the 27 leaderboard rows.

#ia#benchmark#seguranca#agentes#claude#metodologia
Is the human era over? Robots faster than Bolt, higher than Sotomayor, 97% made in China — what the Beijing videos prove, and the two that are fake
August 24, 2026

Is the human era over? Robots faster than Bolt, higher than Sotomayor, 97% made in China — what the Beijing videos prove, and the two that are fake

In 2009, Usain Bolt ran 100 m in 9.58 s and nobody came close for 17 years. On Saturday, 22 August, in Beijing, a robot covered the distance in 9.39 s in an official heat — and a year ago the winner of the same event ran 21.50 s. I read the 14 most shared posts of the week of the World Humanoid Robot Games 2026, downloaded the 11 videos, pulled them apart frame by frame, read the Chinese press, the fact-checkers and the rulebook. The answer has two halves. The first is yes: what happened in Beijing is real, it is bigger than 2025 by an order of magnitude — 2,056 robots, 666 teams, 16 countries, five days — and it is more impressive than the captions say. The second is what the captions hide: the 9.32 s Elon Musk reposted is not the number on the scoreboard (18 s into the video it reads 9.39, heat 9, and the robot that won is not the Honor one); the two most frightening clips on the timeline are fake; and the right question for each video is not "what time did it run" — it is "who was in control". The robots first, then the ruler. Updated on 26 Aug: the final closed the games at 8.64 s.

#robotica#humanoides#china#ia#fact-check#video
August 14, 2026

Claude's text watermark is not a hidden stamp — it is the way the words get picked

The phrase "watermark" makes almost everyone picture the wrong thing: a hidden stamp, an invisible character, something added to the text. It is none of that — nothing is inserted, and what changes is the source of the randomness when the model picks between two words that would serve equally well. I explain the mechanism from scratch, first without a single technical word (the way home and a secret number agreed with a friend), then with the names that appear in the documentation. Every limit falls out of the mechanism for free: why short text barely marks, why code barely marks, why proofreading your own text barely marks, and why the mark will never say who wrote it. I also separate what the press will merge over the coming days: a watermark in text and a C2PA credential in a file are different mechanisms. Seven original figures, made in code. And the caveat that changes how you can use this: the API that would let anyone check a text does not exist yet — the article treats that as a promise, not a fact.

#ia#claude#anthropic#marca-dagua#regulacao#didatico
August 14, 2026

How to know if an AI model runs on your computer — and why

The question from anyone who wants to run a language model on their own computer is always the same: does this fit here? The answer has two parts, and only one of them grows while you talk. I explain both from scratch, three times in a row: first without any technical word, then with the names that appear in the documentation, then with the formula and the real numbers from Qwen3.8-27B, released this month — 27,781,427,952 parameters, 64 layers of which only 16 hold cache that grows, 15.82 GiB of weights in Q4_K_M and 16.14 GiB of cache at maximum context of 262,144 tokens. Along the way, three traps that round numbers hide: K is 1,024 and not a thousand, GB is not GiB, and the model name is rounded. Every number in this article comes from a program published alongside it: any reader can redo the math on their own machine.

#ia#llm#hardware#quantizacao#didatico
August 14, 2026

Anthropic Raised the Risk Score on Its Own Models: What the 186-Page August Report Says

For the first time Anthropic lowered its own grade — and reclassified the past along with it. I read all 186 pages of the August 2026 risk report: what matters is not the score but the five process failures the company describes, including a training-data contamination that hit nearly every Claude model with a knowledge cutoff after December 2024 and was discovered after the report's own coverage date.

#anthropic#seguranca-de-ia#claude#governanca
GLM-5.3: Z.ai delayed the weights citing cyber capability and published 2,436 vulnerabilities as proof — 2,239 of them never left discovery
August 14, 2026

GLM-5.3: Z.ai delayed the weights citing cyber capability and published 2,436 vulnerabilities as proof — 2,239 of them never left discovery

Z.ai delayed GLM-5.3’s open weights claiming the model’s offensive capability grew faster than expected, and offered as proof a public ledger of 2,436 vulnerabilities in real software. I downloaded and counted the whole ledger: 92% of the findings were never reported to anyone, exactly one is marked as sent to a maintainer, no embargo deadline is declared anywhere, and not a single record credits a model — 21% of them used Claude Code as the harness.

#glm#z.ai#llm#benchmarks#open-weights#seguranca
August 11, 2026

Claude Code: the statistic in circulation is 4,000 times smaller than the one anyone can measure

The most complete statistics page about Claude Code credits its download numbers to an npm package that never existed: 404 in the registry, zero captures in the Wayback Machine. The real package has a public API, no key and no paywall — 429 million cumulative downloads, 44.4 million in the very month the page published "111,000+". I measured every number that could be measured, and the scorecard closed at four right, four wrong and two with no source that allows a verdict: the survey of "15,000 developers" had 906 respondents; the "22,000 stars" were 71,847 on 1 March; the "open source" has a LICENSE.md of fourteen words, three of which are "all rights reserved". On the other side, the least believable claim on the list — 4% of all public GitHub commits — checks out as an estimate, and Anthropic itself echoed it in the Series G announcement. The detail that organises the scorecard: the errors all point the same way, downwards. The product measurable by public API is larger than the product described by the page that existed to promote it. And the primary data nobody uses yields the Brazilian angle: 26.0% of Claude.ai usage in Brazil comes from computer and mathematical occupations — above the global average (23.8%) and the United States (21.1%).

#ia#claude-code#anthropic#estatisticas#fact-check#brasil
August 10, 2026

Brazil ranks 5th in the world in Claude usage — and that means less than it seems

Brazil is the fifth country in the world in Claude usage — and unlike almost everything in circulation, the number comes from a primary source: the microdata Anthropic itself publishes under CC-BY. I downloaded the dataset and measured. The same measurement says what the headline does not carry: in per-capita intensity we are 61st of 121 — we sit at the top because we are big, not because we are intense. Volume is not intensity, and the money version of that confusion (run rate is not revenue) props up nearly every "Claude statistics" page of 2026. I checked the most complete of them source by source: the US$ 47 billion run rate is a strong month annualised; the CFO, under oath, declared more than US$ 5 billion since founding in the same window as the public "~19 billion" — not a smoking gun, a speedometer versus an odometer. And the only number anyone can reproduce for free gets one line on the page; the leak gets the headline.

#ia#claude#anthropic#estatisticas#fact-check#brasil
August 3, 2026

ChatGPT: the 900 million came from a funding announcement — and the ruler is in another document

The two numbers holding up any "ChatGPT statistics" page — 900 million weekly active users and 50 million subscribers — come from the same sentence of a US$ 110 billion funding announcement. I went and read the post: the definition of "weekly active user" is not there. It exists, but it lives in another document, published four months later, with a narrower scope. And the denominator has a traceable paternity: the COO measured 400m against total population, OpenAI's paper with the NBER measured 700m against the adult population, and the page that gave me this story went back to total — "10% became 11%" looks like progress and is a switched ruler. On the same ruler it would be 11.3% and 14.5%. Plus the Brazil step: Portuguese is ChatGPT's second non-English language, and two outlets described the same OpenAI briefing with different metrics in the opening line.

#ia#chatgpt#openai#estatisticas#fact-check#brasil
August 2, 2026

How many people use AI? I checked the giant numbers — they count accounts, not people

2.42 billion people use generative AI, says the headline — but the source itself warns against reading that number as people. I checked every step of the pyramid against primary sources, added the step no international version has (Brazil), and the result changes the reading: the world’s paying users are ~1% of humanity, and the coding-agents bubble is 0.14%. You are not behind — the feed is the bubble talking to itself.

#ia#adocao#estatisticas#fact-check#brasil
Noisy-TV in LLM agents: I swept four literatures for the trap — nobody has formalized it, and I measured it in my own agent
August 2, 2026

Noisy-TV in LLM agents: I swept four literatures for the trap — nobody has formalized it, and I measured it in my own agent

In 2018, a curiosity-driven agent froze, hypnotized, in front of a TV tuned to static — the noisy-TV problem, which reinforcement learning spent seven years taming. I swept four literatures for the same trap in the curiosity instrumentation of LLM agents: nobody has formalized it. And it is not hypothetical — in the experimental agent I keep on my own machine, the instrument's noise (σ=0.177) is larger than the signal it is supposed to measure.

#agentes#llm#curiosidade#noisy-tv#embeddings#reinforcement-learning
August 2, 2026

AI statistics in 2026: the scoreboard, checked number by number

Data centers "will consume more than 1,000 TWh in 2026 — the equivalent of Japan," the AI statistics roundups repeat. I went and checked: the IEA itself has retired the number; the current series measures 485 TWh. I checked the most-circulated AI statistics one by one, at the primary source, and the scoreboard has checks-out, half-truth, fossil, and zombie — with one pattern: the omitted denominator ("53% of the population" was the US, ages 18-64). Plus the Brazil step no international version has: 84% of university students have already used genAI, while the country sits outside the top 15 for private investment.

#ia#estatisticas#fact-check#energia#brasil
V4-Flash-0731: DeepSeek published the table where it loses 9 to 0 — and that is the best piece of the launch
August 1, 2026

V4-Flash-0731: DeepSeek published the table where it loses 9 to 0 — and that is the best piece of the launch

The headline says DeepSeek’s cheap model beats the house flagship. The model card says more: Opus 4.8 wins all nine rows of the table — published by DeepSeek itself. Why announcing your own defeat works when you cost 89 times less, what the jump from 7.3 to 54.4 with no new architecture says about post-training, and what the first 48 hours outside the harness confirmed and debunked.

#deepseek#llm#api#benchmarks#open-weights