Categories
Articles Videos

AI / LLM Context Caching: Mechanics, Economics, Storage and Data-Retention Risk

How prompt caching works, why cached input is billed at a fraction of the input price, where the cache physically lives, and what it means for data retention. 2 October 2026.

Executive summary

  1. What is cached is not text — it is internal model state. The provider stores the attention key/value (KV) tensors computed during prefill for a prompt prefix and reuses them when a later request begins with that identical prefix. The raw prompt does not need to live in any shared text store for caching to work.
  2. Cached input is cheap because a cache hit converts a compute-bound job into a memory-read job. Recomputing the KV state for a long prefix is the dominant cost of serving that prompt; reusing it costs a fraction of that. Cached reads have settled around 10% of the full input price across major providers, with 1.25–2× write premiums where cache creation is charged, and best-effort discounts where it is free.
  3. The cache lives on the provider’s serving infrastructure, short-lived and tiered: GPU/HBM first, spilling to GPU-local storage for extended retention (up to 24 h on OpenAI and Azure), host RAM or disk tiers in distributed setups, and — uniquely — durable distributed disk arrays on DeepSeek, cleared within hours to days. It is scoped to the customer (organisation/workspace/project/account) and is never directly customer-accessible.
  4. Data-retention risk is real but bounded, and heavily provider-specific: cached state is derived customer data retained for minutes up to 24 h (or days on DeepSeek), expiry is not immediate erasure, isolation is organisation-level rather than per-user-level, and cache-hit timing side channels are demonstrated in academic research. Customers with strict zero-data-retention requirements face specific model-level conflicts that must be checked explicitly (Section 5).
Categories
Articles

The $2 Trillion Question About Anthropic

Anthropic at $2 Trillion: Right About AI, Wrong About the Price?

The technology is real. The question is who captures the value — and Buffett answered this one about cars and planes a century ago.

Anthropic is heading for a public listing at a reported $2 trillion.

I want to start with what I actually believe, because it matters.

This technology is real, and it will unlock an enormous amount of economic value. Anthropic went from $386m of revenue in 2024 to $4.59bn in 2025 — twelve-fold growth. Q2 2026 revenue hit $11.5bn, with two consecutive profitable quarters and 300,000+ business customers.

Anyone calling AI “all hype” isn’t paying attention.

So the interesting question isn’t whether AI creates value. It’s who captures it — and whether the price leaves anything for you.

The most under-discussed fact in this debate is Anthropic itself

Anthropic was founded in 2021. Within roughly four years it went from nothing to overtaking OpenAI — the company that invented this market, with the multi-year head start, the Microsoft partnership, and the ChatGPT brand.

On enterprise LLM API spend, Anthropic now leads with about 40% share. OpenAI has slipped to the high-20s.

Impressive. But read it again as an investor:

If a newcomer can reach parity with the leader in five years, the leader did not have a moat.

That is what a moat is supposed to mean — protection against exactly this. We have instead watched model leadership change hands repeatedly between OpenAI, Anthropic, Google, DeepSeek and xAI, with benchmark leads lasting months, not years.

The cost structure says the same thing. Anthropic spent $7.33bn on compute in 2025 — 1.6× its entire revenue — and has committed roughly $518bn to infrastructure over the next decade, about 80% reportedly non-cancelable.

That is usually described as a moat. I’d call it the opposite: the price of staying in the game, which every serious competitor is also paying. Capital intensity that everyone must match doesn’t protect returns — it raises the stakes and converts flexible costs into fixed ones owed even if pricing collapses.

Categories
Articles

Scaling AI at Controlled Cost: What the Numbers Say

Everyone is benchmarking which model is smartest.

From a finance seat, the more useful question is: what does one unit of work cost — and does that cost hold as you scale?

I’ve spent months running my own AI models and applications, pushing hundreds of millions of tokens through them. The real story turned out to be the economics.

Per million tokens (output), top-tier frontier models cost up to ~$50.

The cheapest open-weight options: $0.60.

That’s up to 80x.

No negotiation or optimisation closes that gap. It’s a different cost structure — and it changes which projects are viable.

Categories
Articles

The Most Valuable Asset On Your Balance (that most people forget)



We talk endlessly about portfolios, property, and net worth. But for most of us, the single largest asset we will ever own never appears on a financial statement.

Your human capital — the net present value of your future earnings.

Here’s the simple fact: if you earn for another 25 years, that’s not an abstract idea — it’s a number you can compute. And at a 5% discount rate, a steady $100k a year for 25 years is worth ~$1.41M today. That’s a real, quantifiable asset sitting at the top of your personal balance sheet, whether you acknowledge it or not.

Categories
Articles

Why, as a finance professional, I use Python instead of Excel



Finance is a lot of process. Specific data, specific steps, specific decision points. And most of that process is rule-based and logic-driven — which means it’s automatable.

Over the years, the share of my analytical work done in Excel has steadily shrunk — in favour of code editors. Today, less than 5% of my analytical work happens in Excel. The other 95%+ runs in Python and purpose-built analytical tools.

What Python gives me that Excel can’t:
– ⚡ Speed — fetch & process data in seconds, not spreadsheets
– ✅ Reliability — same logic runs the same way, every time
– 🧠 Complexity — models scale cleanly beyond a grid of cells
– 🔄 Flexibility — adjust & update in code, not a rebuild

The workflow: Fetch → Process → Structure → Act — each step automatable.

Excel is still great at what it’s great at. But for the process of finance, Python is the better tool. I’m a realist: I use both. The more complex and repeatable the work, the more I reach for code.

Categories
Articles

I hired a (personal) AI employee.

I hired a (personal) AI employee. It works for me in the background, and it’s now a permanent part of my “staff”.

For over a month I’ve run Hermes Agent — an autonomous AI that lives on my hardware, has its own memory, and does things. Not a chatbot. A worker.

What it actually does for me:

🎬 1,500+ videos & articles → downloaded, transcribed, summarized, and filed into a searchable learning library. I just drop a link; it does the rest. And it structures the whole process — turning a chaotic stream of random finds into a clear curriculum and learning plan I can actually work through. A very different outcome from what the social media algorithm will feed you.

🔎 A research assistant that knows my standards. I throw topics at it and it scours various sources — with references, context, and analysis. It’s learned what matters to me in research, so the output comes back the way I’d want it, not generic.

📚 5 Anki (language learning) collections (66,901 media files) synced to my phone via a self-hosted server it deployed — and debugged when the SQLite schema broke.

🔐 My VPN — added a TCP fallback, managed the firewall, and diagnosed a carrier-NAT issue that was blocking me on mobile.

🧠 A private digital coach — its own persona, memory, and Telegram chat, checking in on my career and health. Runs on local models on my own hardware. My conversations never leave my infrastructure.

📝 A notes repository it maintains and cross-references automatically.

What I learned about managing it:

  1. Autonomy is the unlock. I brief it, it executes end-to-end, and reports back. I don’t babysit it.
  2. It fails — and self-corrects. Crashed containers, locked databases, retries. Resilience beats perfection.
  3. Memory compounds. It remembers my personal targets and goals, my health habits, my taste in music. Every interaction gets smarter.
  4. Privacy is a feature. My most personal conversations — with my coach — run on local models. No cloud, no training on my data.
  5. You still need judgment. It’s a tool, not a replacement for thinking. I review the important stuff.

The honest take: I’m a realist — I know what technology can and can’t do, and I don’t chase hype. But AI is maturing, and the potential is very significant. After weeks of real use, Hermes Agent has earned its place as permanent infrastructure, not a novelty. It’s the difference between having an assistant and having a staff.

Curious about hiring your own agent — and your own private coach — on your own hardware? Happy to share what I’ve learned.

Categories
Videos

Exponential growth and epidemics

A brilliant video explaining the math involved in the exponential growth of epidemics and how measures to reduce exposure can reduce the spread by orders of magnitude. Video credit: 3Blue1Brown

Categories
Videos

What is the Fourier Transform? A visual introduction.

Credit for the video: 3Blue1Brown

Categories
Videos

Introduction to Monte Carlo methods

Credit for the video: Alon Honig

Categories
Videos

Space: The Next Trillion Dollar Industry