4 min read

Monday Memo

Study the beats, not just the prose | AI shopping hype| What is 'compute?'
Monday Memo

A Byte of Knowledge

Pull up the ads spending the most right now and read the copy.

Plenty of it is mediocre, the "3 p.m. crash," the phrasing every chatbot grabs.

And yet it keeps winning, because the structure underneath does the heavy lifting.

Look closer and every line has a job: one names the enemy, the next hands over the mechanism, and by the close the reader feels late to something everyone knows.

Which is why swiping just the words gets the wrong layer.

Instead, ask each line of a winner what it's there to do, then write the answers down in order.

Now you're holding the beats.

Hand those to your AI, with each one's job, and it rebuilds the sequence around your product.

Hand it only the words, though, and you get flat copy back without the skeleton that made it work.


🚀 WATCH: How These AI Copy Bots Are Producing World-Class Sales Copy 50X Faster Than Even The "BEST" Copywriters On The Market…

(Plus… They Don't Get Sick, Miss Deadlines, Or Ask For Raises Either!)

Watch the full AI Copywriting Tell-All Video Here


Something to Ponder

We Asked – AI Answered

Our Question: 

"In the world of AI what all constitutes ‘compute.’ Breakdown the full ecosystem and the key factors that make up compute.”

AI's Answer: "Compute" Decoded - What Actually Powers AI

🧮 What Compute Is

An AI model is a giant grid of numbers called parameters. Every time it reads or writes a token (a word or word fragment), it multiplies those numbers against your input. Compute is the capacity to do that math, and the industry counts it in FLOPs (floating-point operations... single calculations on decimal numbers).

A model spends compute in two places:

  • Training: the model processes trillions of tokens and tunes its parameters, a bill the lab pays once.
  • Inference: the finished model answers questions, a bill that runs every time anyone uses it.

🔢 How the Math Adds Up

Researchers use a simple rule of thumb:

Training FLOPs ≈ 6 × parameters × training tokens

Meta's Llama 3 shows it working: 6 × 405 billion parameters × 15.6 trillion tokens ≈ 3.8 × 10²⁵ FLOPs, the same figure Meta reported.

To feel the scale: if all 8 billion people on Earth each did one multiplication per second, nonstop, they'd need about 150 million years to match that run. Meta did it with roughly 16,000 Nvidia H100 chips.

Inference runs on a smaller rule, about 2 × parameters per token, so each token that model writes costs roughly 800 billion FLOPs.

⚙️ Why AI Runs on GPUs

CPU: a handful of brilliant accountants solving complex problems one after another

GPU: thousands of clerks each doing one simple multiplication at the same moment

AI math is mostly simple multiplication, so GPUs dominate. Chipmakers also shrink the numbers: moving from 32-bit to 8-bit or 4-bit precision packs far more math into each second, and Nvidia's Blackwell chips added native 4-bit support.



📊 Four Factors That Set Real-World Compute

These decide how much of a chip's advertised speed you actually get:

  1. Raw speed: Nvidia rates its H100 at nearly 1,000 trillion FLOPs per second on 16-bit AI math.
  2. Memory: High-bandwidth memory (HBM) sits beside the chip and feeds it parameters. When data arrives too slowly, the chip sits idle... engineers call it the "memory wall."
  3. Networking: Training splits across thousands of chips that constantly swap results, and slow links stall the whole run.
  4. Power and cooling: One Nvidia GB200 NVL72 rack draws about 120 kilowatts and needs liquid cooling. U.S. data centers used 4.4% of the nation's electricity in 2023, per Lawrence Berkeley National Lab.

Reality check: Meta's Llama 3 paper reported hitting 38-43% of its chips' theoretical peak. The rest went to waiting and coordination.

🏭 The Supply Chain Underneath

ASML (Netherlands) → sole maker of the EUV lithography machines that print leading-edge chips

TSMC (Taiwan) → manufactures most cutting-edge AI chips on those machines

SK Hynix, Samsung, Micron → supply the HBM

Cloud giants → assemble everything in power-hungry data centers

Few links have a quick substitute, and the most critical one sits on an island Beijing claims as its own. TSMC's Arizona fab made its first Nvidia Blackwell wafer in October 2025, an early step toward domestic supply.

💼 What It Means for Your Business

Every AI tool you pay for resells this math, usually priced per token. Bigger models and longer answers burn more FLOPs per task, so two choices drive your bill:

  • Model size: Run routine jobs like tagging leads or summarizing calls on small, cheap models, and save frontier models for strategy and high-stakes copy.
  • Output length: Reasoning models "think" in extra tokens before answering, and you pay for those tokens too, so reserve them for problems that genuinely need deep reasoning.

Thanks for reading the Monday Memo.

Until next time!

The AI Marketers

P.S. Help shape the future of this newsletter – take a short 2-minute survey so we can deliver even better AI marketing insights, prompts, and tools.

[Take Survey Here]