What We’re Reading (Week Ending 27 September 2026)

The best articles we’ve read in recent times on a wide range of topics, including investing, business, and the world in general.

We’ve constantly been sharing a list of our recent reads in our weekly emails for The Good Investors.

Do subscribe for our weekly updates through the orange box in the blog (it’s on the side if you’re using a computer, and all the way at the bottom if you’re using mobile) – it’s free!

But since our readership-audience for The Good Investors is wider than our subscriber base, we think sharing the reading list regularly on the blog itself can benefit even more people. The articles we share touch on a wide range of topics, including investing, business, and the world in general. 

Here are the articles for the week ending 27 September 2026:

1. The Most Important Market in AI is the Middle – Tomasz Tunguz

Yesterday, Anthropic released a new model & cut its price. Ninety minutes later, OpenAI did the same.

Most business AI use is the messy middle: multi-step workflows that need a smart enough model at a price a company can afford. It is the most important part of the market today, & it is where the competition is fiercest. The price cuts are the evidence.

In June, Anthropic set the frontier price at $10 & $50 per million tokens with Fable 5. In July, OpenAI answered with GPT-5.6 Sol at $5 & $30, matching that capability at a third of the cost per task.

The Opus line had never moved. Opus 4.5, 4, 4.8 & 5 all listed at $5 & $25 per million tokens. Yesterday’s cut was the first…

…But the right tail of the market is thinner than almost anyone forecast. Anthropic’s Fable 5.1, its most capable & most expensive model, commanded only 3.7% of gateway spending in its first twelve days. Its predecessor peaked at 13.2% when access was restored in July, then fell to 4.9% a month later when Opus 5 shipped at half the price. Among large corporate accounts, frontier models fell from 53% of token consumption in early August to 45% by September…

…As intelligence per dollar explodes, the distribution of tokens may shift to commodity. Whether that happens will determine the economics of the AI market.

2. History Doesn’t Repeat, but It Rhymes – Venky Ganesan

Between 1998 – 2000, teams of engineers left Cisco, Nortel, Lucent and Bell Labs to build optical switches, terabit routers and long-haul DWDM gear, and they were bought, sometimes before the product existed, for billions of dollars in stock. Cerent went to Cisco for $6.9 billion. Xros, a Menlo portfolio company from before my time at the firm, went to Nortel for $3.25 billion with 90 employees and no shipping product. Qtera, also $3.25 billion, same story. Chromatis went to Lucent for $4.5 billion. Siara went to Redback for $4.3 billion while, in the words of one reporter, still in the prototyping phase.

Here is the cast, in case the analogy isn’t obvious. Cisco was the Nvidia of its day: the picks-and-shovels company that everyone had to buy from, briefly the most valuable company on earth at around $555 billion and 200 times earnings. Nortel, Lucent, JDSU and Ciena were the hyperscalers: the giants with enormous balance sheets and inflated stock who bought the startups and built the capacity. The CLECs, WorldCom, Qwest and Global Crossing were the customers, the ones actually placing the orders, and they were doing it with about a trillion dollars of borrowed money. And the optical startups were the neo labs…

…Thirteen startups were bought pre-revenue by strategics in 1999 and 2000. Between them they had raised roughly $500 million of venture capital. They sold for about $31 billion. Every investor in every one of those companies made money, and some of the multiples were absurd: Xros returned something like 130 times the capital that went in. Now look at what happened to the products. Ten of the thirteen were killed within about two years…

…The venture returns had nothing to do with whether the product worked. They had to do with whether somebody with an inflated stock price bought you before the music stopped.

Now look at the ones that didn’t sell early. Thirty-two companies raised $100 million or more privately between 1999 and 2002, about $5.9 billion in total. Twenty-two of them, 69 percent, returned roughly nothing. Caspian raised $317 million and shut its doors in 2006. Procket raised $272 million and sold its intellectual property to Cisco for $89 million. Pluris raised $215 million and never shipped a router. Three returned capital  back. Two, Amber and Flarion, returned three –  four times. Five got public before the window closed. Corvis went out in July 2000 at a $27.6 billion valuation with zero revenue, raised $1.1 billion, and only ever deployed about six switches. Avici, CoSine, Tellium and ONI followed the same path. The public got the loss instead of the VCs, and all five are gone…

…Vintage determined who won. Founded in 1996 or 1997 and exited by the middle of 2000, you did great, product or no product. Founded in 1999 or 2000 and still private in 2001, you got nothing, team or no team. Nobody could see the line while they were standing on it. In the summer of 2000 the second group looked exactly as smart as the first. The only difference was the calendar.

The end, when it came, was fast. Ciena agreed to buy Cyras for $2.6 billion in December 2000. The deal closed in March 2001 at $1.1 billion. The price fell 58 percent between the handshake and the check. Cisco’s revenue was growing 60 percent through December 2000; by the April 2001 quarter it was down 25 percent and the company wrote off $2.25 billion of inventory…

…The parts that rhyme are the structure. There is a picks-and-shovels monopoly at the center with a handful of customers who are each a huge share of its revenue. There is a layer of giants spending more than they earn to build capacity for demand that hasn’t shown up yet: the big four hyperscalers have guided to somewhere between $720 and $760 billion of capex this year, Alphabet’s capex exceeded its operating cash flow last quarter for the first time in its history, and the industry has issued about $344 billion of AI-related debt this year alone. There is a customer layer, the labs, whose spending commitments dwarf their revenue: OpenAI at roughly $40 billion of run-rate against around $1.4 trillion of compute commitments. And there is a startup layer priced off the last round rather than off anything built…

…The parts that don’t repeat matter too, and I don’t want to pretend they don’t. Nvidia earns money. Cisco was at 200 times trailing earnings; Nvidia is at 29x, with a 62 percent net margin. The customers this time have real revenue growing very fast, not just borrowed money; the CLECs never had anything like $40 billion of run-rate, let alone Anthropic’s. The hyperscalers are, for now, cash-generative franchises with businesses that exist independent of AI, which Nortel and Lucent were not. If you had to bet on which layer of this stack survives a downturn, the answer today is much better than it was in 2000…

…You cannot pick your way out of this. The 2000 telecom investors were good at picking; they picked Larry Roberts and Tony Li. What they didn’t do was size for the possibility that the calendar, not the team, would decide the outcome…

…And who were some of those people who funded those companies? They were VC legends like Vinod Khosla, Pierre Lamond, Paul Ferri, Ed Anderson,  and Jim Breyer.

3. The AI Inference Revolution Is Here – Matthew S. Smith

An untrained LLM is like a jumble of Scrabble tiles on a table. Instead of single letters, though, the tiles show fragments of words, called tokens. Everything you’d need to write almost anything is present, but nothing makes sense.

Training a model organizes this jumble using a guessing game played at scale. The model is shown real text with the next token hidden and asked to predict what comes next. After each guess, the correct token is revealed and then compared to the prediction, and the difference is used to calculate the model’s accuracy. The game is played not with a single sentence but over billions of passages.

While a real game of Scrabble can be played over a bag of chips and a few drinks, AI training is computationally intense. The model updates its parameters through backpropagation, a process that repeatedly calculates how each of a model’s billions or trillions of parameters should shift to make the next prediction better. This is why tech giants are building larger data centers than ever before.

Eventually the model’s creator decides further training isn’t worth the cost, and the guessing game stops. Backpropagation ends, the parameters are frozen, and the LLM becomes a pretrained model. Fine-tuning—a short training run on smaller, more specialized data—adds final tweaks, and the model is deployed.

Next comes inference. This is the process of using the deployed model, which, now that it’s been trained, has learned to spit out Scrabble tiles—tokens—in a sensible order.

You might think that AI inference is less computationally demanding because the backpropagation calculations used to update parameters are eliminated. But Sudeep Bhoja, founder and CTO of the inference-hardware company d-Matrix, explains that inference adds new challenges.

The models are “autoregressive” in nature. That is, the next output depends on the previous one. “So to generate the next token, you have to read all of the weights and all of the [context] from the previous token,” explains Bhoja. The context includes all of your prompts, all of the LLM’s replies, and all of the files you upload. It’s a lot of data and a lot of processing.

An LLM generates its reply in two phases: prefill and decode. Prefill is the model reading a prompt. It processes every token at once, computing how each token relates to all the others. This operation is called attention, and it’s a defining characteristic of the transformer architecture behind modern LLMs. It allows them to respond to a word in its sentence, paragraph, and larger context rather than on its own. Think of it like arranging Scrabble tiles before you place them in a game. Many players move tiles around to imagine how they connect. Self-attention plays a similar role, though instead of moving physical tiles, each token sends a query to the others and receives a score indicating the token’s relevance.

These queries result in two types of vectors: the keys and values. They are typically placed in a store called the KV cache. This isn’t strictly required, as a model could instead recompute these vectors with each new token it generates. But nearly all LLMs use a KV cache to reduce how much computing they do. The KV cache is stored in memory and becomes a scratchpad to which the LLM can return to understand a conversation, and though it starts small, it can swell to dozens of gigabytes.

Prefill is a problem that can be easily divided up and worked on in parallel. This is why GPUs became the dominant AI accelerator as LLMs surged in popularity. Graphics rasterization (computing the color of every pixel on a screen) is also massively parallel, so GPU architectures were a natural fit.

Next comes decode. Here, the model generates its reply one token at a time. At each step it takes the most recent token, weighs it against everything in the KV cache, uses that information to predict the next token, and adds the new token’s key and value to the cache. Then it repeats in sequence, token by token.

This is where the autoregressive nature of the model works against inference speed. Predicting each token requires reading the entire model from memory, and that model consists of possibly tens to hundreds of gigabytes of parameters (the numbers representing what the model learned in training). Crucially, this is in addition to the memory required to store the KV cache.

As a result, the movement of all this data through memory often requires more bandwidth than inference hardware has available. So at least some of the computing parts of a GPU sit idle as it waits for data…

…The big players—Nvidia and Amazon—are going for an all-chips-on-deck approach. Nvidia’s GPUs and Amazon’s Trainium training accelerators are still great for part of the inference workload: the prefill stage, where all the context keys and values are calculated. But to accelerate decode, the part where new tokens are generated, they are looking to new, memory-centric architectures from smaller players.

In Nvidia’s case, the smaller player was Groq (not to be confused with Grok, the family of LLMs trained by SpaceXAI). Nvidia purchased intellectual property and hired talent from Groq at the end of 2025, and just three months later at the Nvidia’s GTC 2026 conference, Jensen Huang unveiled the Nvidia Groq 3 language-processing unit (LPU). Groq’s architecture relies on memory—in its case, SRAM—built directly into the chip’s architecture…

…Amazon Web Services (AWS), for its part, struck a deal with Cerebras, to pair the Trainium accelerator with Cerebras’s Wafer-Scale Engine 3 (WSE-3). Cerebras takes a similar approach to Groq, though at a much larger scale. WSE-3 turns an entire silicon wafer into a single chip that contains over 4 trillion transistors. The design doesn’t connect to external memory but instead etches 44 gigabytes of SRAM into each wafer. “We store the [model] weights on the SRAM,” says James Wang, formerly director of product marketing at Cerebras who has since moved to SpaceXAI. “So that’s easily 40 to up to 80 billion parameters that we can support on one chip.”

Amazon plans to use AWS Trainium chips for prefill, and Cerebras for decode. But Cerebras’s chips can also go it alone in inference. WSE-3 was deployed by OpenAI to power GPT-5.3-Codex-Spark, a variant of the company’s coding mode, outputting over 1,000 tokens per second. For comparison, OpenAI’s standard GPT-5.4 deployment outputs 50 to 125 tokens per second…

…Most computers store numbers in a 32-bit or 64-bit format. These determine how many bits are available to represent a single number. If too few bits are available, the number can’t be stored without losing information. The quality of an LLM benefits from more-precise number formats, but this creates a problem for inference performance. More-precise numbers aren’t free. The bits that describe them take up more space in memory and require more silicon and energy to compute…

…The process of converting an LLM from a more-precise number format to a less-precise format is called quantization, and it’s been in use for several years. However, researchers are finding new ways to quantize models down while retaining a large majority of the model’s quality.

Nvidia recently created a new 4-bit number format, NVFP4, for this purpose. AMD, Intel, and Qualcomm have instead rallied around a competing 4-bit number format called MXFP4 that Nvidia also contributed to developing. “It’s the black art of AI,” says Buck, of Nvidia. When Nvidia quantized DeepSeek-R1 from FP8 to NVFP4, scores on seven major benchmarks degraded by less than one percent while performance improved by three times, the company says.

4. Frontier Overhangs – Ben Thompson

One of my go-to examples in Anthropic’s Safety Superpower was the company’s decision to predicate Fable usage on Anthropic holding onto all customer data for at least a month; this was a big deal, and I argued at the time that Anthropic was making a bet that its models were good enough to convince enterprises to give up on zero data retention:

It’s pretty significant, I think, that Anthropic is declaring that not retaining data is no longer an option, at least if you want access to their best models. Yes, today, that retention is for safety purposes only, and not for training; it’s plausible, however, that Anthropic’s lead becomes so significant that they quietly announce that they are going to train on that data as well, and companies will feel they have no choice but to go along. That additional training data, of course, will only further increase Anthropic’s lead, and all of this will be justified because Anthropic has already clearly decided they are the only ones who can be trusted to be in charge.

In fact, Fable wasn’t good enough: customers pushed back, and Fable usage stayed relatively low; when Fable 5.1 was released, the Anthropic-gets-to-keep-your-data provision was gone. This is evidence of Christensen’s theory in action: customers demonstrated the willingness to base their model-choice decision on something other than pure performance, namely, data retention policies.

This doesn’t, in and of itself, suggest that new model capabilities aren’t desired; it does, however, suggest that current model capabilities are “good enough” for customers to not do whatever is necessary to get access to the cutting edge, which reduces the value of the cutting edge to its proprietors, and gives credence to the strategy of Microsoft and others focused on separating harness and model. Pure capability no longer translates directly into a moat…

…The key for the frontier labs, then, is to build those user touchpoints while they have superior capabilities. However, this is where Meta’s recent launch of Muse is a bearish signal. Muse is, by a significant margin, the best and most approachable personal agent product I have tried. Meta deserves a tremendous amount of credit for the product work they have put in, as well as the massive infrastructure commitment entailed in providing users with a very capable virtual machine for free. Oh, and of course they deserve credit for the Muse Spark model undergirding Muse.

That noted, Muse Spark 1.3, the most advanced Meta model, is still not state-of-the-art, and that is the bearish signal: it is good enough for a very good personal agent product, and critically, a personal agent is much stickier than a chatbot. Once you have put all of your information into a personal agent and actually incorporated it into your day-to-day life, it is much more of a challenge to change to something else. This is in contrast to Codex/ChatGPT and Claude Code: yes, you may have developed your own set of skills and understanding of how each harness works, but at the end of the day the relevant artifacts (i.e. your code) are in GitHub, and it’s not that much of a lift to point a different agent and harness at those artifacts if the alternative is better and/or cheaper (or doesn’t want to keep all of your data).

In short, model capability is good enough that compelling products — products that actually have moats — can now be built, and from a business perspective it would do Anthropic and OpenAI good to devote more of their resources to actually building such products.

5. The US Dollar’s Two Competing Forces – Joe Weisenthal and Tracy Alloway

But Englander sees this line of thinking as a dead end. He writes:

Think of the US as an asset manager or hedge fund that converts low-yielding foreign savings into higher-returning assets. The USD’s value partly reflects how much of that higher return the US must surrender to attract the capital it needs, and partly the rest of the world’s confidence that the US can deliver these returns. Investors may also allocate capital to the US if it is seen as a better preserver of wealth than elsewhere. In each case, the USD’s value is part of the price the rest of the world pays for these services.

Under this framework, a stronger USD reflects higher expected returns from US asset management or greater demand for US custody and safe-haven services – hence the ‘dollar smile’ (referring to USD strength during both strong growth and rising safe-haven demand). A weak USD signals the opposite, or an expectation that policy or financial market uncertainty will make these services available at a lower cost in the future.

Now what’s interesting is not that the dollar has fallen since President Trump’s inauguration, but that the periods of pronounced weakness have tended to coincide with what many people perceive to be “policy missteps” such as Liberation Day, Scott Bessent’s bond buyback announcement, and the row over Greenland…

…Now unfortunately, as Englander notes, the weaker USD hasn’t “bought” the US anything in terms of trade imbalances. Real non-oil imports continue to go basically straight up, and the US hasn’t seen any particular boom in manufactured exports.

To put it another way, if there is a “good” version of dollar weakness, we’re not getting it.

Disclosure: We currently have a vested interest in Alphabet, Amazon, Mastercard, Meta Platforms, Microsoft, and Visa. Holdings are subject to change at any time. 


Disclaimer: The Good Investors is the personal investing blog of two simple guys who are passionate about educating Singaporeans about stock market investing. By using this Site, you specifically agree that none of the information provided constitutes financial, investment, or other professional advice. It is only intended to provide education. Speak with a professional before making important decisions about your money, your professional life, or even your personal life. We currently have a vested interest in Alphabet, Amazon, Meta Platforms, and Microsoft. Holdings are subject to change at any time. 

Leave a Reply

Your email address will not be published. Required fields are marked *