<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0">
  <channel>
    <title>Umbrelog Engineering Notes</title>
    <link>https://umbrelog.com/blog</link>
    <description>Lessons on incident investigation, production debugging, and reducing time-to-understand.</description>
    <language>en-us</language>
    <item>
      <title>Root Cause Analysis (RCA): A Practical Guide for Production Incidents</title>
      <link>https://umbrelog.com/blog/root-cause-analysis-production-incidents</link>
      <guid isPermaLink="true">https://umbrelog.com/blog/root-cause-analysis-production-incidents</guid>
      <description>Root cause analysis (RCA) for production incidents: incident timelines, evidence vs assumptions, postmortem structure, common mistakes, and a practical incident response workflow — with an RCA example.</description>
      <category>Incident response</category>
    </item>
    <item>
      <title>How to Investigate Production Incidents (Step by Step)</title>
      <link>https://umbrelog.com/blog/how-to-investigate-production-incidents</link>
      <guid isPermaLink="true">https://umbrelog.com/blog/how-to-investigate-production-incidents</guid>
      <description>How to investigate production incidents step by step: confirm the alert, find an anchor error, build a timeline, check deploys and infra pressure, communicate findings, and write the postmortem.</description>
      <category>Incident response</category>
    </item>
    <item>
      <title>The Hidden Cost of &quot;What Deployed?&quot;</title>
      <link>https://umbrelog.com/blog/hidden-cost-what-deployed</link>
      <guid isPermaLink="true">https://umbrelog.com/blog/hidden-cost-what-deployed</guid>
      <description>Every incident eventually asks what shipped. The answer is rarely in one place — and that delay costs more than teams admit.</description>
      <category>Engineering</category>
    </item>
    <item>
      <title>The Same Error 5,000 Times Is Still One Incident</title>
      <link>https://umbrelog.com/blog/same-error-five-thousand-times-one-incident</link>
      <guid isPermaLink="true">https://umbrelog.com/blog/same-error-five-thousand-times-one-incident</guid>
      <description>Why duplicate errors overwhelm on-call engineers without creating understanding.</description>
      <category>Engineering</category>
    </item>
    <item>
      <title>MTTR Is a Lagging Metric. Nobody Measures Time-to-Understand.</title>
      <link>https://umbrelog.com/blog/mttr-lagging-metric-time-to-understand</link>
      <guid isPermaLink="true">https://umbrelog.com/blog/mttr-lagging-metric-time-to-understand</guid>
      <description>When the graph turns green but the Zoom is still open — why fast recovery and finished understanding are not the same clock.</description>
      <category>Engineering</category>
    </item>
    <item>
      <title>We Centralized Our Logs. Incidents Didn&apos;t Get Faster.</title>
      <link>https://umbrelog.com/blog/centralized-logs-incidents-not-faster</link>
      <guid isPermaLink="true">https://umbrelog.com/blog/centralized-logs-incidents-not-faster</guid>
      <description>Centralizing logs solves collection — not the five-tab assembly every Sev-2 still runs by hand across GitHub, APM, and Slack.</description>
      <category>Engineering</category>
    </item>
    <item>
      <title>What Building a Logging Platform Taught Me About Incident Investigations</title>
      <link>https://umbrelog.com/blog/incident-investigations-lessons</link>
      <guid isPermaLink="true">https://umbrelog.com/blog/incident-investigations-lessons</guid>
      <description>Why teams still spend more time understanding incidents than fixing them — and what building Umbrelog taught us about the gap between logs and understanding.</description>
      <category>Engineering</category>
    </item>
  </channel>
</rss>
