10 min read

Need to Know News - September 4th, 2026

Hershey's AI finds where to stack s'mores. | Google's AI video edits like a pro. | Nearly 3 in 4 shoppers now let AI pick what they see.
Need to Know News - September 4th, 2026

In this week's Need to Know News edition:

πŸ€– GPT-6 Astra is officially here...is AGI too?

πŸ€– Runway showed an AI that builds the interface as you click through it... and one of its first target uses is your storefront.

πŸ€– Leaked Meta files show an agent that keeps working after you close the app... and one line in them hints where it might show up next.

And a whole lot more!


GPT-6 Astra Has Officially Landed

OpenAI released the model it held back for weeks of safety work, and the headline number is hard to believe. On ARC-AGI-3, the puzzle test built to stump machines, GPT-6 Astra scored 99.9% where its predecessor managed 7.8%.

The ARC Prize Foundation says it effectively reached human parity. OpenAI calls Astra its most intelligent and aligned model ever, and the practical claim is simpler: this one operates your computer for you better than ever.

πŸ“‹ The Details: Astra fills out web forms, updates CRM records, tidies your calendar, and drafts summaries inside your email or docs, then builds a slide deck that follows your template rather than inventing its own.

Through Sites in ChatGPT it can create, host, and share a website from one prompt. On OSWorld, the computer-use test, it finished tasks in roughly 40 minutes each against 75 for GPT-5.6 Sol, and on Agents' Last Exam it beat Claude Opus 5 using about 65% fewer output tokens... which matters, since you pay per token.

🎯 Why You Need to Know: The assistant that answered questions is becoming one that does the task itself, inside the software you already use, and the lower token use means delegating routine work costs less at the same time.

Two cautions from OpenAI's own notes: Astra crosses the company's Critical cybersecurity threshold, so it refuses exploit-building tasks and runs under misalignment monitoring that can pause legitimate work, and its written reasoning is harder to monitor than Sol's.

⚑ Your Move: As always with new models, test it with your best prompts and see the real world difference.

Full Story

Altman Apologizes as Paying Subscribers Wait on Astra

At 7:49 PM on September 3, Sam Altman announced the best model in the world. By 7:53 his replies were full of paying subscribers asking where it was, and at 1:09 AM he typed "i am hopeful that you can use it this weekend! but can't promise yet." The next morning, he called the rollout messy and apologized.

πŸ“‹ The Details: Access went to a limited set of organizations first, leaving paying subscribers with a coming-days promise. Codex lead Thibault Sottiaux posted the remedy at 11:12 PM: every day a paid ChatGPT plan goes without Astra earns one banked reset, a spare usage-limit refresh held for later. The staged rollout was planned... OpenAI described it September 1 after pausing frontier training over the Hugging Face incident.

🎯 Why You Need to Know: A messy launch, by OpenAI's own word, still gives you compensation calculated per day, and those resets accrue whether anyone on your team notices.

πŸ“‘ Watch For: Whether Pro accounts, first in line, get Astra over the September 5 and 6 weekend Altman hoped for. If it slips past Monday, the resets keep accruing.

Full Story


πŸš€ WATCH: How These AI Copy Bots Are Producing World-Class Sales Copy 50X Faster Than Even The "BEST" Copywriters On The Market…

(Plus… They Don't Get Sick, Miss Deadlines, Or Ask For Raises Either!)

Watch the full AI Copywriting Tell-All Video Here


Snap Ads Now Bid Against the Install Data You Already Trust

App advertisers have always had two sets of numbers: the platform's own, and the mobile measurement partner (MMP) counting installs across every channel, which rarely agree. Snapchat's Unified Attribution brings the MMP data into Snap Ads Manager so its bidding engine works from the figures your finance team believes, and it's now live worldwide.

πŸ“‹ The Details: It works today with AppsFlyer and Adjust. A real-money gaming advertiser saw a 26% lower cost per acquisition than its SKAN-based campaigns, and Mohegan Sun Online Casino posted 89.8% higher return on ad spend and 77% lower cost per install by week four. Those are Snap's own case studies.

🎯 Why You Need to Know: AppsFlyer's partner lead put it plainly: for the first time, Snap's bidding engine and your measurement read from the same data. Until now you bid on one set of numbers and reported on another.

⚑ Your Move: If you run app installs through AppsFlyer or Adjust, ask your Snap rep to switch Unified Attribution on for one campaign and run an otherwise identical SKAN-based campaign beside it for four weeks.

Full Story

Williams-Sonoma's Personalized Visits Bring 9x the Revenue

Ask Otto, Pottery Barn's new AI assistant, for a rug and it'll size the room, match it to the sofa, and hand you to a human designer when needed. Williams-Sonoma rolled Otto out across Pottery Barn in August, and last week's earnings call gave numbers.

πŸ“‹ The Details: Over 70% of Otto's conversations resolve without a human, says tech chief Sameer Hassan, and shoppers who use Olive, its older sibling, convert at three times the rate of those who don't. The bigger number: a visit where the company tailors the experience now generates roughly nine times the revenue of an average one, up from two times last year.

🎯 Why You Need to Know: CEO Laura Alber says the tactile parts of the business can't be replaced by AI, and the results support that... the assistant narrows choices and hands off, and revenue rises where it hands off to a person at the right moment.

⚑ Your Move: Pull two numbers on your chatbot: the share of conversations that end without a human, and conversion for people who used it versus skipped it. If you can't get both, start there.

Full Story

Gemini 3.8 Flash Lands in Sheets and Search at 3.7's Price

Three weeks after 3.7 Flash, Google is back with 3.8, its third Flash in six weeks, at the same price. Google calls it its best reasoning and coding model yet. The catch is that the price is temporary.

Source: Google

πŸ“‹ The Details: The $0.75 per million input tokens and $3.75 output is introductory and doubles after December 31, 2026. Until then it's live for Google AI Pro and Ultra subscribers in the Gemini app, AI Mode in Search, and Gemini in Sheets. Google says 3.8 works harder, running extra reasoning steps and using more tokens at higher effort. A second variant, Flash Cyber, stays behind the new Fairwind Program for vetted defenders.

🎯 Why You Need to Know: If you're on Pro or Ultra, the model inside your spreadsheet improved at no extra cost, and Google says it got much better at resisting prompt injection, where hidden text in a file hijacks the AI.

⚑ Your Move: Open Gemini in Sheets and give it the multi-step analysis you'd normally break into five prompts, like a channel-by-channel margin read.

Full Story

JD Sports Found Relevant Offers Beat Bigger Discounts

JD Sports ran a structured test of its loyalty offers and came away with a finding that should worry anyone whose retention plan is a bigger coupon. Working with Leeds-based AI start-up, HyperFinity, the retailer found a more relevant offer consistently beat a more valuable one.

πŸ“‹ The Details: The engine tested buy-again category offers, cash rewards, spend multipliers, and cross-sells. Heavy spenders responded best to multipliers, while lighter spenders reacted about the same to cash and multipliers, so the extra money on that group was mostly wasted. In-store loyalty sales rose over the trial, though neither company will say by how much, and JD now describes a multi-million-pound opportunity in the first year of scaling it.

🎯 Why You Need to Know: This is a budget finding as much as a marketing one... if a well-chosen small offer outperforms a generous generic one, you can spend less on loyalty offers while sales rise.

⚑ Your Move: Split your next loyalty send. Half get your usual discount, half get a smaller offer chosen from what each person bought last, and compare redemption and basket size.

Full Story

Leaked Files Show Meta's Hatch Agent Working While the App's Closed

Close the app and Meta's Hatch keeps going. That's the premise of an unreleased personal agent that, per pre-release materials TestingCatalog reviewed, fills out forms, orders dinner, books restaurants, and runs deep research on your behalf.

πŸ“‹ The Details: Onboarding runs three steps... connect accounts (email, calendar, Instagram, OpenTable), pick starter tasks, then name your agent and pick its personality. Web, iOS, and Android builds exist, the iPhone one already on its fourth test version, and Meta has considered a premium tier as high as $199.99 a month. The most notable detail is a reference to Messenger and WhatsApp Companions, which would put an autonomous agent inside apps billions already open.

🎯 Why You Need to Know: If an agent that can shop lives inside the app your customers already message in, discovery, ordering, and reservations could happen in WhatsApp with no site visit. All of this comes from leaked files, so features may shift or vanish before launch.

πŸ“‘ Watch For: A waitlist or invite-code page from Meta, and whether the first version ships with browser control only. TestingCatalog pegs September as plausible.

Full Story

Optimizely Sells Marketers a Chief of Staff You Can Hire

81% of B2B marketing leaders switch between two or more disconnected AI tools every week, per Optimizely's own survey of 2,000-plus. Its answer is Virtual Teammates: AI coworkers with a job title, a permission set, and a memory, so nobody re-explains the company to a fresh chat window.

πŸ“‹ The Details: The launch roster covers a Chief of Staff that preps meetings, an SEO and AI Search Analyst that audits where you show up in AI answers, and a CRO Manager that reads your experiments. You hire them from a directory, adjust duties in plain language, and put them on a schedule. Each acts under its own identity with its own permissions and audit trail, and you decide where a human signs off.

🎯 Why You Need to Know: The memory is the useful part: a teammate that retains last quarter's decisions does the follow-up your current chatbot forgets the moment you close it. The 81% comes from Optimizely, which sells the product.

πŸ“‘ Watch For: Whether Optimizely publishes pricing and a general availability date, and how the SEO and AI Search Analyst reports on AI visibility.

Full Story

Runway's Solaris Generates the Website While You're Using It

Every app you've ever used was finished before you opened it, and unchanged until the next update. Runway's Solaris removes the code step. It generates the interface itself, frame by frame, as you click and drag, so the picture on screen is the software.

πŸ“‹ The Details: Runway calls it an Interface World Model, treating clicks like a text prompt. The demos show a clothing store where you drag a shirt onto a photo of yourself, and a storefront that rearranges itself for each visitor. In a 250-person study, people preferred Solaris over a Claude Opus 5-coded version 71% to 21% for natural behavior, though Runway admits legible text is still a weak spot, a problem for anything with a price tag. It's early access only.

🎯 Why You Need to Know: Runway lists a storefront that redraws itself for each shopper as a target use, so this is aimed at your website, not only at games.

πŸ“‘ Watch For: Which partners Runway names for the public launch. A retailer on that list means real storefronts, and a demo with readable prices means the text problem is fixed.

Full Story



Anthropic's Commerce Agent Kit Claims 35% Bigger Carts

Carts up to 35% larger and shoppers 60% more likely to finish buying... that's what Anthropic says retailers running shopping agents on Claude have seen, and this week it published a kit: working agents, guardrails, and a Claude Code plugin meant to get one running in days, timed for holiday planning.

Source: Anthropic

πŸ“‹ The Details: The shopping agent sits on your site, takes a request like a tent, sleeping bag, and stove, and builds the cart. The merchant agent handles inventory and promotions, flagging an item about to sell out before a promotion starts, though a person approves every change. It's available today, and Wix says its engineers had an agent taking prompts within fifteen minutes.

🎯 Why You Need to Know: The conversion numbers are Anthropic's, so treat them as a sales claim. Still, the guardrails matter: catalog-pinned prices and a no-manipulation rule are what a shopping bot needs before you'd trust it with a customer.

⚑ Your Move: Skip the repo and open the self-guided demo at claude.com/solutions/commerce. The retail example tells you whether this fits your catalog before any engineer touches it.

Full Story

ChatGPT Ads Hit a $1 Billion Run Rate in 200 Days

Google's second full year of selling search ads brought in about $440 million, and at the time that was considered fast. OpenAI passed it in about seven months. ChatGPT's ad business crossed $1 billion in annualized revenue Monday, roughly 200 days after its first ads, and OpenAI is guiding to $2.5 billion booked this year.

πŸ“‹ The Details: Ads run in more than 40 countries on the free and Go tiers, and the self-serve platform opened Monday to advertisers in India, Europe, the Middle East, and North Africa. The same day, the European Commission designated ChatGPT a Very Large Online Search Engine on 159.1 million monthly EU users, reported under legal obligation, with compliance due by December.

🎯 Why You Need to Know: Forbes contributor Jon Markman expects the near-term fight to be over budgets that were never in search, the brand money now in social feeds and retail media... your money. Nothing in Monday's numbers shows spend leaving Google.

⚑ Your Move: If you buy in a region that opened Monday, open a self-serve account and run a small test against your best social ad.

Full Story


Thanks for reading.

Until next time!

The AI Marketers

P.S. Help shape the future of this newsletter – take a short 2-minute survey so we can deliver even better AI marketing insights, prompts, and tools.

[Take Survey Here]