Press "Enter" to skip to content

Posts published in “Software and Frameworks”

Hands-on notes, model benchmarks, open-source libraries, developer APIs, and agentic workflows that you can actually deploy today.

Siri AI, a profoundly more capable and personal assistant powered by the next generation of Apple Intelligence, is here  

Apple;

Siri AI is a completely reimagined version of Siri that is more personal and powerful. With detailed responses and rich, natural conversations, Siri AI helps users get more done than ever before.

Built on Apple Intelligence, this new version of Siri draws on personal context understanding to help users find what they need in the moment across messages, emails, photos, and more. And with even more systemwide app actions, Siri AI can help users with tasks such as drafting an email from scratch or editing and sharing a set of photos.

This is the Siri Apple always intended it to be.

In order to use Apple Intelligence, you need to be using iOS 27, iPadOS 27, macOS 27, watchOS 27, and visionOS 27 is available on iPhone 16 models or later, iPhone 15 Pro, iPhone 15 Pro Max, iPad mini (A17 Pro), iPad models with M1 or later, MacBook Neo (A18 Pro), Mac models with M1 or later, Apple Vision Pro, Apple Watch Series 9 or later, Apple Watch Ultra 2 or later, and Apple Watch SE 3 when paired with an Apple Intelligence-enabled iPhone nearby.

Google Announces Gemini 3.8 Live and 3.8 Live Extended Thinking

Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, marking a major leap in real-time, audio-to-audio artificial intelligence. These models are built specifically to handle fluid, back-and-forth voice interactions while simultaneously processing complex underlying logic.

Core Capabilities

  • Gemini 3.8 Live: Optimized for low-latency, natural dialogue. It allows users to speak organically with the AI, supporting real-time interruptions, dynamic tone adjustment, and fluid conversational flow without typical voice-assistant delays.
  • Gemini 3.8 Live Extended Thinking: Tailored for complex tasks that require higher background reasoning. By leveraging parallel processing, the model executes multi-step logic and problem-solving “behind the scenes” while maintaining an uninterrupted, natural spoken conversation.

Key Benchmarks & Features

  • Audio Intelligence: Achieves 97.7% on Big Bench Audio, establishing a new benchmark for spoken instruction comprehension and context retention.
  • Continuous Audio Stream: Operates natively in an audio-to-audio framework rather than converting voice to text and back, drastically reducing response latency.
  • Parallel Reasoning: Extended Thinking actively resolves edge cases, coding problems, or computational questions mid-conversation without pausing the verbal output.

These updates transform voice agents from simple command-execution tools into active collaborative partners. Developers can deploy Gemini 3.8 Live for live customer support, conversational tutoring, or real-time translation. Meanwhile, the Extended Thinking variant enables complex hands-free workflows—such as debugging code via voice or working through multi-variable strategy problems during a call—making voice interaction far more versatile across technical and professional fields.

You can read the full text of Google’s announcement at Blog.Google.

Google announces Gemini App for Windows 10 and 11

Google has expanded its desktop lineup with the launch of the dedicated Gemini app for Windows. Designed as a system-level utility, the client provides instant AI assistance directly over active applications on Windows 10 and 11.

Google added in their announcement 3 ways of using the Gemini App in your PC;

1. Access Gemini instantly with a keyboard shortcut.
Press Alt + Space on your PC at any time to open Gemini over your active work. Whether you need a quick fact-check on a document or a few catchy title ideas for a presentation, you can get the help you need and jump right back into your workflow.

2. Power through deep work in a dedicated workspace.
Inside the new app, you can access everything you use Gemini for already. Hand off multi-step tasks to Gemini Spark, your 24/7 personal AI agent, or ask Gemini to draft a project summary by pulling information directly from your Google apps like Gmail and Google Drive.

3. Create images and videos.
Bring creative concepts to life directly from your desktop. You can generate custom images with Nano Banana for a presentation, direct a high-quality video with Gemini Omni, and best of all you can do it all in one convenient place.

Google added that they developed the gemini app to be lightweight and quiet in the background without degrading PC performance.

To download the Gemini App for Windows just visit: gemini.google/desktop.

Google stated that this release represents the initial desktop footprint for Windows, with additional native OS capabilities planned for future rollouts.

OpenAI Releases ChatGPT Images 2.5: Faster Speeds, Targeted Editing, and Creative Tools

OpenAI has launched ChatGPT Images 2.5 (powered by the GPT-Image-2.5 model family) that will be available in all ChatGPT, ChatGPT Work, and Codex users across desktop, mobile, and web. Rather than just offering a baseline image generation update, this release centers on reducing friction in creative workflows through faster generation times, stronger reference consistency, and intuitive visual editing features.

In their official announcement, OpenAI said “Images 2.5 produces more natural lighting and richer textures, is better at preserving the subjects in your reference photos, and follows editing instructions more reliably across multiple turns. We’ve also reduced image generation latency by up to 50% compared with Images 2.0, so you can generate images and refine your concepts more quickly.”

The company also introduces a new features in ChatGPT;

Sketch is a new feature that lets you draw directly in ChatGPT as a reference for your final image. Templates make it easier to start creating across some of the most popular image formats, like flyers and product photos. You can also now place comments directly on images for more focused editing, and you can share prompts you’ve used to let others try your ideas with their own photos and details.

For developers, we’re introducing two new models in the API. GPT‑Image‑2.5 Flare brings the same improvements in quality, editing, and speed, and GPT‑Image‑2.5 Sunburst offers an extra level of precision for detailed creative work with longer generation times.

Pricing and availability

Model Modality Input Cached input Output
gpt-image-2.5-sunburst Image $8.00 $2.00 $30.00
Text $5.00 $1.25
gpt-image-2.5-flare Image $8.00 $2.00 $30.00
Text $5.00 $1.25

You can go here for additional pricing info.

OpenAI Unveils GPT-6 Astra, Its Most Advanced Model Yet

OpenAI has launched GPT-6 Astra, which it’s calling the world’s most intelligent and aligned model. The company says Astra brings together years of research and big bets across pre-training, reinforcement learning, and alignment, and that it is state-of-the-art on computer use, browsing, software engineering, cybersecurity, science, and professional work.

Technical Capabilities & Benchmark Highlights
On native computer use and workflow automation, Astra interacts directly with digital environments by browsing websites, filling out CRM records, managing calendars, and running frontend QA checks. On the OSWorld 2.0 latency simulation, Astra scored 72.6% in roughly 40 minutes per task, outperforming GPT-5.6 Sol (65.7% in 75 minutes) while completing tasks 47% faster.

Regarding context limits, the model features a 1,050,000-token input context window and supports up to 128,000 output tokens, allowing it to process massive repositories, long documentation sets, and complex codebases without breaking context.

Designed as OpenAI’s primary software engineering model, Astra reached 57.9% on Terminal-Bench 4.0 for advanced coding and system engineering. It retains long-horizon conversation context within Codex, allowing it to search past interactions asynchronously to solve complex infrastructure bugs.

For scientific and mathematical reasoning on graduate-level evaluations, Astra recorded 96% on GPQA Diamond, 98% on FrontierMath Tier 4, and 99.9% on ARC-AGI-3.

In cybersecurity capabilities, Astra achieved a 100% score on ExploitBench (without production safeguards) and 42.4% on ExploitGym for vulnerability analysis and defense patching. Production versions include strict guardrails that refuse automated exploit creation.

Benchmark Comparison

Evaluation Metric GPT-6 Astra GPT-5.6 Sol Claude Fable 5.1
Agentic Coding (Terminal-Bench 4.0) 57.9% 37.3% 55.8%
Computer Use (OSWorld 2.0) 72.6% (~40 mins) 65.7% (~75 mins)
Cybersecurity Vulnerability (ExploitBench) 100% 78.5%
Scientific Reasoning (GPQA Diamond) 96%

API Pricing and Deployment Rollout
Standard API pricing for Astra is set at $10 per million input tokens and $50 per million output tokens. The model is deployable across the OpenAI API, OpenRouter, Microsoft Azure, and AWS Bedrock.

OpenAI is rolling out access across ChatGPT Plus, Pro, Business, and Enterprise accounts. Workspace access is toggled off by default for enterprise administrators to allow security auditing, and eligible enterprise API customers can enable Zero Data Retention (ZDR) policies upon deployment.

GPT-6 Astra marks an operational pivot for OpenAI. Rather than simply retrieving answers or writing code snippets, the model is engineered to execute long-horizon, multi-step workflows across terminal windows, local applications, and enterprise browsers with lower token consumption.

Meta Launches Muse Spark 1.3 with 1M Token Context and Enhanced Reasoning

Meta has officially unveiled Muse Spark 1.3, the latest flagship model from Meta Superintelligence Labs designed for long-horizon coding and autonomous agentic workflows.

Operating with a 1 million token context window, the model introduces key architecture updates engineered to sustain extended multi-step execution within a single thread. Rather than acting as an overeager automation tool, Muse Spark 1.3 is designed for collaborative reliability: it actively asks clarifying questions when prompts are ambiguous, flags structural uncertainties, and requires user confirmation before taking consequential or irreversible actions. Benchmarks indicate that the model completes complex engineering tasks using roughly 20% fewer tool calls and 25% fewer tokens compared to Muse Spark 1.2. On evaluations like DeepSWE 1.1 (75.4%) and long-context retrieval (98.5%), it posts top-tier performance while competing directly alongside industry peers.

The model specifically targets enterprise software engineers, AI agent developers, and cost-conscious tech teams. By focusing on low tool-call overhead and multi-turn stability, Meta aims to solve the key bottlenecks facing developers building complex refactoring systems, autonomous web agents, and CI/CD tools.

Available immediately via the Meta Model API and integrated into Muse Code, Meta offers standard pricing at $1.25 per million input tokens and $4.25 per million output tokens, along with a discounted $0.10/$0.20 Contributor tier, delivering high-tier reasoning at aggressive operational efficiency.

Google Announces Gemini 3.8 Flash and 3.8 Flash Cyber

Google has officially launched Gemini 3.8 Flash alongside a dedicated security variant, Gemini 3.8 Flash Cyber. Marking Google’s third Flash release in six weeks, the updated generation focuses on software engineering, multi-step reasoning, and long-horizon agentic workflows.

The standard Gemini 3.8 Flash delivers substantial benchmark improvements—including leading scores on coding evaluations like DeepSWE v1.1—matching or outperforming larger frontier models at a fraction of the cost. Developers can fine-tune “effort levels” per task, allowing the model to perform deeper, iterative reasoning loops for complex enterprise problems.

Meanwhile, Gemini 3.8 Flash Cyber is engineered specifically for defensive security teams. Offered to verified defenders via Google’s Fairwind Program, it excels at autonomous vulnerability discovery and patch generation across complex codebases.

Both models are now available through Google AI Studio, Vertex AI, and the Gemini API.

Anthropic announces Claude Fable 5.1 and Claude Mythos 5.1

Anthropic said in their announcement, “Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards. Fable 5.1 is generally available, while Mythos 5.1 is available only through our trusted access programs; its safeguards are specifically designed to support work in cybersecurity and the life sciences. Alongside its increased capabilities, Fable 5.1 takes important steps towards addressing the feedback we’ve received from customers on price, data retention, and safeguards.”

May be related to Anthropic’s previous announcement, “Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%.”

Key Upgrades & Technical Performance

  • Next-Level Coding & Benchmarks: Fable 5.1 sets new performance records, scoring 52.6% on the Terminal-Bench-Science 0.1 benchmark (compared to 24.7% for Fable 5) and reaching 73.4% on CursorBench 3.2.0. The model excels at finding root causes in complex codebases—in one instance identifying a rare bug in an external vendor library that had eluded human engineers for years.
  • Mythos 5.1 for Specialized Research: Available strictly through trusted access programs, Mythos 5.1 utilizes specialized, high-tier safeguards designed for advanced cybersecurity vulnerability discovery and scientific research in the life sciences.
  • Significant Cost Reduction: Typical token-based workloads will see an estimated 25% cost reduction compared to Fable 5 due to lower pricing on cache reads. Highly agentic workflows can yield total savings of up to 45%.
  • Enterprise Frontier Safeguards (EFS): Anthropic introduced EFS to deliver zero data retention capabilities by allowing enterprise users to store data directly within their own cloud infrastructure rather than on Anthropic servers.

Benchmark Comparison

Metric / Benchmark Claude Fable 5.1 Claude Fable 5 GPT-5.6 Sol
Agentic Coding (Terminal-Bench 4.0) 55.8% (60.9% for Mythos) 42.0% 37.3%
Scientific Research (Terminal-Bench-Science 0.1) 52.6% 24.7% 22.4%
Multidisciplinary Reasoning (Humanity’s Last Exam – with tools) 65.0% 63.8%

ICYMI: Google Launches AI-Powered Shopping Features in Gemini for the Philippines

Google has expanded its e-commerce capabilities by rolling out AI-driven shopping features within the Gemini app in the Philippines. The update allows local users to research, compare products, check prices, and find where to purchase items—all inside a single conversational chat.

Google announced the new shopping feature in Gemini late last June.

Anthropic to Adjust Claude Code Limits: 25% Permanent Increase Replaces 50% Promo on Sept. 14

On the official Claude Dev X/Twitter account, the company announced: “Starting September 14, we’re permanently raising standard weekly limits in Claude Code by 25% for Pro, Max, Team, and seat-based Enterprise plans. Until then, the current 50% increase will remain in place.”

At first glance, a 25% permanent increase sounds like good news. But that 25% is measured against Claude Code’s original baseline — not against the temporary 50% boost users have been enjoying since May. Once the promotional period ends, the effective allowance drops.

The company added that “Compared to today, this works out to a 17% reduction in weekly limits on Claude Code. We’re working on exciting changes that will make it feel like you’re getting more from Claude, while having more visibility and control of your usage. Can’t wait to share them.”

Users and commentators quickly called out the mismatch between the celebratory framing and the actual outcome, with many noting it’s a familiar pattern as AI providers manage rising compute costs while trying to keep subscribers happy.