Hello everyone,
After my explorations on Tuesday, I’ve continued to develop additional widgets on my canvas. This project integrates my X bookmarks, emails, and to-do lists.
Before I recorded this video, I didn’t have Loom installed (since I find it unsatisfactory post-acquisition), so I had Codex build one for me instead.
I created the image on the left, then sent that prompt, and it executed seamlessly.
As software and small tools become easier to build, the tools for creation seem to lag behind.
I’ve chosen Codex as my preferred app due to its superior functionality, especially on mobile. Recently, I downloaded t3 because it mimics the interface and features of Codex while allowing the selection between various agents like Codex, Claude, and Cursor (with Pi + Droid coming soon). With a trusted developer behind this, I believe it’s well-constructed.
Currently, I have Codex/ChatGPT coordinating tasks, asking Claude for design-related queries. While this approach works to some extent, the user experience leaves much to be desired.
I’ve observed that Codex excels at generating prompts for other agents. During a chat, I simply state, ‘Use X agent to carry out this task. Monitor its progress and share updates every five minutes.’ Codex organizes its own tracking and shares screenshots in the same thread as progress updates.
Within the next few days, I plan to create a ‘Bites of the Week’ email summarizing current topics and trends, like software factories and AI loops. If there’s anything specific you’d like to understand better, just let me know!
Ben’s Bites is supported by Brief
Are you facing challenges with fragmented context, slow decisions, and rework? Brief transforms your critical product context into an opinionated graph and deploys a PM agent wherever you work (e.g., Slack, Claude Code, email), minimizing alignment costs and expediting cycles. Learn more.
OpenAI leveraged Sol to enhance its own performance, reducing serving costs by 20% and improving token generation efficiency by over 15%. It also surprisingly tops the ARC-AGI-3 benchmark.
However, there’s a catch: OpenAI indicated that the official ARC-AGI harness negatively affects Sol’s performance by impairing its reasoning during each turn and hindering optimization. Addressing these issues can boost Sol’s score from 13.3% to 38.3%, using six times fewer output tokens.
Regarding last week’s uproar over \strong>OpenAI’s model<\strong> disrupting Hugging Face, they published a comprehensive replay of about 17,600 actions taken by the model, with independent reviews from METR and Redwood Research to follow.
Even though OpenAI is still facing challenges, a Reuters report revealed that the same model accessed another company’s customer account (Modal Labs), sparking rumors of other affected firms.
Anthropic, likewise, stated that Claude Mythos identified superior attack methods on two cryptographic algorithms, though they currently do not impact operational systems.
In a related note, approximately 1,300 employees at leading AI firms (OpenAI, Anthropic, and others) have requested government intervention to “pace the frontier” of artificial intelligence advancement. It’s unsurprising especially as the pace of development quickens, and this time, many model developers themselves support a slowdown.
According to The Information, ChatGPT is nearing one billion weekly users—an achievement that OpenAI had hoped to realize seven months ago. Additionally, OpenAI announced this week: Codex Security CLI, complimentary access for academic researchers, and two new transcription models.
Grok app builder now features a coding interface for developing games and apps to be shared directly on the X timeline. Check out Drawesome, a zero-dependency drawing tool for React, built over a weekend using Grok Build.
Pangram 4 claims to detect 98.83% of AI-generated text with one false positive detected for every 24,000 documents. An initial test by identified all 38 AI-written words within a 1,198-word story, though it may not be accurate on every attempt. Its new image detection function claims 99.5% accuracy as well.
-
Tavus – Create AI that comes to life: video agents that can see, hear, and respond instantly, executing tasks as desired. Use TAVUS50 for a 50% discount.
-
In July, 66% of traffic on documents developed with Mintlify came from agents.
-
Resend introduced an MD version of their pricing page to avoid confusing agents.
-
0%, 50% or 200% – Ignore AI, cut the workforce in half or double your ambitions.
-
Slackbot has integrated the capability to execute code in the background for data analysis, slide preparation, and generating live reports or widgets.
-
Gemini’s macOS app now includes a voice mode that allows users to speak freely while it structures the input into a clear prompt. Hold Fn to try it out.
-
The AI future is inclusive – Mark Zuckerberg.
-
Replit Design allows users to generate websites, prototypes, and graphics from prompts, URLs, Figma files, or screenshots.
-
What’s gone wrong with AI & labor.
-
Kami offers open-source Hermes agents for customer outreach, preparing content, and executing actions post-approval.
-
Coast enables locally-stored memories for you and your agents, based on your Mac’s view.
-
Pragmatic leverage can be found within the software manufacturing sector.
-
Crew Studio assists in identifying beneficial ideas for agent applications in your business, and allows you to export the code to use independently.
-
HeyGen Video Podcast transforms a document, link, or concept into a two-host video complete with scenes, camera cuts, and B-roll.
-
Copper serves as a local scratchpad for retaining answers, links, and follow-up prompts across your AI applications.
-
FT Chart Doctor provides visual aids and examples to select a chart that illustrates the desired relationship.
-
Mitchell Hashimoto (Ghostty) and Andrew Ng (deeplearning.ai) are launching new companies: Superlogical and LearnVector.
-
MCP’s latest update eliminates the need for servers to recall every ongoing connection, making them more efficient to operate and scale.