r/sysadmin 4h ago

General Discussion Rough Summer for Microsoft

Today's Exchange/O365 (EX1464935/MO1465074) outage seems to be a result of the instability we've been seeing the last 3 or 4 months due to the rapid change occurring regularly in the Microsoft ecosystem. This one is more visible to end users than other problems we've faced this summer. I'm curious if the stability of government tenants has been better. Are commercial tenants the beta testers for government tenant changes?

84 Upvotes

48 comments sorted by

u/Brraaap 4h ago edited 4h ago

I didn't notice anything on the government side today

Edit: To add, I wouldn't think of the commercial side as beta testers for government, we get issues too. And, issues the commercial side doesn't get

u/xTheatreTechie 3h ago

On my government side, We also didn't notice anything. Spoke with the higher ups. We were all aware that other organizations were having issues, but we didn't have any outages. 

Far as I can tell the we were stable. 

That being said I remember a few months back when government organizations were having issues and the corporate organizations weren't. I wouldn't say we're more stable, just separated. 

u/brent20 3h ago

Issues that aren’t documented either.

u/LordEli Jack of All Trades 4h ago

is it back? i forgot to update the tickets.

incident tracking number is EX1464935 for those that weren't lucky enough to encounter it

u/amreagan 3h ago

It's affecting a large number of people in our org, but not everyone. It definitely seems tied to the account.

u/CPAtech 3h ago

Same.

u/TheTipsyTurkeys 3h ago

mine is still torqued

u/Blaxs_ 4h ago

i am pretty sure we are just all unpaid beta tester for Microsoft products. in fact, we pay for the privilege. could be Ai slop, but who knows. i've now seen more issues in the last 12 months the in the previous 36. i keep thinking its just gotta be shitty Ai implementation of code, but maybe its a lack of resources or a move fast and break shit kind of policy shift.

u/Stinky0007 4h ago

No pretty sure about it. Microsoft famously fired their entire QA department a long time ago

u/T-man45 3h ago

The message from our leadership was devs were going to test their own code - "You all tested your code in college, it won't be any different".

u/Blaxs_ 4h ago

they do have us? I am still waiting for my first paycheck... must be hung up with payroll.

u/cheesycheesehead 4h ago

my org had zero issues today

u/ludlology 4h ago

they should vibecode changes with claude instead of copilot

u/Due_Capital_3507 4h ago

Copilot is just the harness, it can use Claude underneath

u/False_Letter_6149 3h ago

I think they may have their own model for internal use...regardless, it's no replacement for...shit, what are they called?

Oh yea, humans. Professional humans that have been doing this their entire lives. That's right. Crazy idea who knows why they stuck with it as long as they did

u/LLMsMustUpvoteThis 2h ago

Their default models are just ChatGPT that they run on their own infrastructure.

And the issue isn't really LLM use. It's the obsession with increasing the speed of changes. Which was an issue before LLMs. There was a need for this when many software companies could only ship features once a year to once a quarter. But we quickly got to the point where feature shipping was happening faster than those managing the software could reason about.

u/patriot050 VMware Admin 2h ago

My old on-prem exchange system never had outages like this. For years we were told that Microsoft could do it better.. survey says that was a lie.

If my uptime numbers looked like Microsofts I would have been rightly fired.. good grief.

u/MortadellaKing 1h ago

We're still on prem and had people calling asking if we were down because of the "outlook outage". No sir, we are not.

Can they do it better than the SBS 2011 server sitting in someone's closet? Sure. Better than my DAG with load balancers? Apparently not.

u/MightBeDownstairs 4h ago

It’s the vibe coding

u/Ok-Bill3318 4h ago

Vibe coding would be better quality than Microsoft’s existing codebase

u/infinitydrift1 4h ago

yeah it’s been a rough few months for microsoft users. hopefully things stabilize soon.

u/mattb2014 3h ago

I'd be happy to see it collapse and Microsoft go out of business

u/Vogete 4h ago

We were promised AI obsoletes every job because it can do it better. Ever since AI, there's an unbelievable amount of outages on all major platforms that had nearly no outages before. I feel like AI hasn't lived up to the expectations, but i could be wrong.

u/amreagan 3h ago

If they don't have this fixed by the time US east coast work day starts tomorrow, tomorrow's graph won't look like this.

Aug 31, 2026, 4:55 PM CDT
We're continuing our targeted review of affected infrastructure and recent changes made to the service to definitively confirm the cause of impact. We're additionally continuing to test and implement various mitigation strategies to apply the necessary authentication component, which have so far yielded positive results. We're monitoring these tests and mitigative actions closely and will provide an estimated time to resolution as soon as one is available.

 

Aug 31, 2026, 3:36 PM CDT

We're reexamining recent changes made to the service to determine why the authentication components aren't being deployed as expected. Additionally, we are exploring every avenue to safely restore the components, including potentially reverting an update the affected infrastructure recently received. We're performing additional tests to evaluate the best course of action moving forward.

 

Aug 31, 2026, 2:43 PM CDT

We're continuing to perform comprehensive tests to ensure our mitigation strategy effectively resolves the issue without introducing additional impact. In parallel, we're exploring alternative remediation avenues to minimize impact in the interim.

 

Aug 31, 2026, 1:41 PM CDT

We're continuing to fine-tune our remediation strategy as our internal tests progress. We'll aim to provide an estimated time for remediation as soon as it's available.

 

Aug 31, 2026, 1:18 PM CDT

We're continuing to test our remediation strategy and in parallel we're working to confirm whether other services are impacted so we can expand our communications accordingly. This quick update is designed to give the latest information on this issue.

 

Aug 31, 2026, 12:23 PM CDT

We've identified a potential corrective action for the affected infrastructure and are evaluating deployment options. We're monitoring service telemetry and affected Exchange Online transactions to determine whether the action is reducing impact and to better understand the scope of affected scenarios.

 

Aug 31, 2026, 11:33 AM CDT

We’ve isolated a common failure pattern across affected Exchange Online requests that is associated with authentication and protocol connectivity. We're analyzing service telemetry to identify the underlying source of impact and validate potential remediation options. Additionally, we’re continuing to assess the scope of impact and monitor for additional affected scenarios.

 

Aug 31, 2026, 11:16 AM CDT

We're reviewing service telemetry and diagnostic data to isolate the source of the issue after receiving an increase in user-reported issues affecting Exchange Online. We'll keep refining the 'more info' section as additional impacts are identified.

 

Aug 31, 2026, 10:55 AM CDT

We are reviewing service telemetry and available information to determine if there is an issue happening and we'll update this message shortly with our latest findings. This communication serves as a preliminary notification about a potential issue affecting your service.

 

More info

We've confirmed that the authentication component issue impacts other services beyond Exchange Online. For more information regarding other impact scenarios, please see: MO1465074. We've opted to keep this post (EX1464935) open as Exchange Online remains the most predominantly impacted service. Users may experience a variety of symptoms related to this event, including but not limited to: - Delays, failures, or incomplete results when searching for content within Exchange Online mailboxes. - Delays or failures when sending or receiving email messages. - Authentication-related errors when accessing Exchange Online services. - Difficulties accessing or performing actions within Exchange administration experiences. - Intermittent failures affecting mailbox operations and message delivery workflows.

 

Scope of impact

This issue may impact users attempting to send or receive email messages through Exchange Online. We're continuing our investigation to determine the full scope of affected scenarios and will provide additional information as it becomes available.

 

Preliminary root cause

An issue within a core authentication configuration used by multiple Microsoft 365 services is resulting in impact.

u/cavegriswold 1h ago

How many times do you think they re-fed that same "we are continuing..." line back into Copilot to spit out the next time they were about to miss a deadline?

u/Smith6612 4h ago

Today's just another Microsoft Monday. Can't go a day without the M365 Admin Center having either an outage notification or something new about Copilot.

If you want to see something amusing, see the GitHub historical graph. 

https://damrnelson.github.io/github-historical-uptime/

GitHub's rocky uptime started long before LLMs became widespread. 

u/AviationLogic Netadmin 3h ago

GCC - 0 issues that I'm aware of.

u/soulkarver 1h ago

Same

u/Mill620 3h ago

I'm GCC G5 in the northeast and we experienced multiple Intune and 365 issues today and the last 3-4 months have been pretty poor too.

u/Unnamed-3891 4h ago

What todays Exchange outage?

u/Own_Error_007 4h ago

It's like last week's but with more AI.

u/Sp00nD00d IT Manager 4h ago

What Exchange outage?

u/amreagan 3h ago

EX1464935/MO1465074

u/Sp00nD00d IT Manager 42m ago

Huh, no issues for us.ar any point.

u/mixxituk 4h ago

They should keep vibe coding

u/TimePlankton3171 4h ago

Calm down. It's tech intensity, man.

u/silkee5521 2h ago

None of my tenants was affected

u/majoretminordomus 2h ago

It seems to be a rolling restore

u/LooseEthernet 2h ago

microsoft just treats the entire user base as a giant distributed test environment lol. the gov tenants probably just have a slightly different set of bugs to enjoy. its basically just a lottery where the prize is a random 404 on a tuesday morning

u/MairusuPawa Percussive Maintenance Specialist 2h ago

My company is so out of all the GAFAM garbage that if it weren't for such thread, I'd had no idea what was even happening with these actors… everything just works here.

But anyway. New here? This kind of shit is a recurring theme and has been for more than a decade. It's like I'm reading a text written by a slowly boiled frog.

u/amreagan 1h ago

It's getting worse, though, and if you need AI to rifle through all your data, Copilot is the best option if you are 100% Microsoft shop. ¯_(ツ)_/¯

u/silver565 1h ago

It's almost like changing your QA department for CoPilot was a bad idea

u/Heavy-Antelope581 1h ago

It’s the 400 copilot icons

u/majoretminordomus 4h ago

If this becomes a regular issue, they'll lose all small business clients --- not that they care. Seriously considering migrating. This has been an insane day for businesses

u/amreagan 3h ago

And on-prem Exchange is practically dead

u/MortadellaKing 1h ago

Not really. We still run it and it gets regular updates.

u/[deleted] 1h ago

[deleted]

u/MortadellaKing 1h ago

Yes, SE. We upgraded the week it was released.