The AI boom is becoming easier to measure—and harder to simplify. Microsoft ended its fiscal year with a $90 billion quarter and rapid Azure growth, while Meta’s own growth came with a striking collapse in quarterly free cash flow. Google pushed generated music toward finer creative control, and two new research papers showed the distance between impressive agent activity and dependable professional results.
Microsoft turns AI demand into a $90 billion quarter
Microsoft closed fiscal 2026 with quarterly revenue of $90.0 billion. Its earnings release put Microsoft Cloud revenue at $59.3 billion, up 27%, while Azure and other cloud services grew 43%. Microsoft 365 Copilot, meanwhile, passed 30 million paid seats.
Those figures offer a rare view of the AI cycle from both sides of Microsoft’s business. Azure reflects demand for the underlying compute and cloud services; Copilot seats reflect businesses paying for AI inside familiar software. The seat count is not the same as usage, retention or disclosed profitability, but it is harder evidence of enterprise adoption than another model demonstration.
Meta’s growth engine meets the cost of the buildout
Meta’s second-quarter revenue reached $60.801 billion, up 28%, but its costs rose 55% to $42.026 billion and net income fell 14% to $15.848 billion. Capital expenditure was $31.08 billion. The sharpest number in Meta’s results was free cash flow: $784 million, down from $8.549 billion a year earlier, a decline of roughly 91%.
One quarter should not be mistaken for a permanent cash-flow rate, and Meta’s cost increase cannot be attributed entirely to AI infrastructure. Free cash flow is also a non-GAAP measure. Still, Meta’s narrowed full-year capex forecast of $130 billion to $145 billion makes the strategic direction unambiguous. Its advertising business is financing one of the largest infrastructure programmes in technology, and the return on that buildout now matters as much as the ambition behind it.
Lyria 3.5 gives AI music finer controls
Google has rolled out Lyria 3.5 in Flow Music, claiming improvements across musicality, lyrics, vocals and creative control. According to the launch announcement, the model can produce richer melodic structures, follow lyric prompts more closely, generate more expressive vocals with better pronunciation, and let users control tempo and duration more easily.
That combination points to the next stage of generative music: fewer novelty clips and more tools that can be directed toward a specific creative result. Google did not publish an independent comparison or listening study with the announcement, so the quality gains remain its own claims. The meaningful test will be whether musicians find the new controls predictable enough to use repeatedly.
Research agents can build the experiment—and still miss the research
A new preprint proposes “shadow evaluations” for open-ended AI research. Frontier agents were given the central questions from two unpublished NeurIPS 2026 submissions, six days of work and thousands of dollars of compute. They completed all the engineering without human assistance but failed to make substantial progress on either research question. The original authors grading the work rejected both outputs.
The study identifies recurring failures in research judgment, creative response, backtracking, resource awareness and instruction-following. A second model and scaffold reproduced the pattern. That makes the result more than a single bad run, but not a final verdict: it is a preprint based on two case studies. Its most useful contribution is the distinction it draws between automating the labour around an experiment and automating the judgment that makes an experiment scientifically valuable.
APEX-Accounting measures the gap between partial credit and reliability
APEX-Accounting places models inside 10 simulated work environments containing accounting systems, spreadsheets, PDFs and other files. Its private evaluation set contains 160 expert-authored tasks covering work such as reconciliation, accruals, transaction posting and reporting.
Claude-Fable-5 leads the reported results with 56.4% Mean Criteria@3. Yet no model exceeds 2.6% on the benchmark’s stricter Pass^8 metric. The contrast is the story: a model can satisfy many rubric criteria while remaining far from dependably completing the whole professional workflow. The headline task set is private and APEX-Accounting is a closed benchmark, so independent researchers cannot fully inspect the evaluation behind those numbers.
The money picture now has a demand side and a cost side. Microsoft’s cloud growth and Copilot seats suggest businesses are buying AI capacity and tools; Meta’s free-cash-flow compression shows how expensive the race can become before those investments mature. Watch for evidence beyond headline adoption: Copilot usage, returns from Meta’s infrastructure, independent Lyria testing, broader research-agent replications and accounting benchmarks that outsiders can audit. Signal, not advice.
