SHAREPLANE AGENT CONTEXT Schema: shareplane-context/2.0 Schema URL: https://next.shareplane.malott.ai/schemas/agent-package.schema.json Projection: Generated public-safe plain text. Not canonical Markdown. Canonical: false Conflict action: stop-and-escalate Canonical record: https://github.com/pinklon/pinklon-shareplane-next/tree/33d227f6000da2491f19209edde916d8846f287f/content/artifacts/the-lost-discipline-of-fault-isolation/artifact.json Canonical record SHA-256: b14b699b1e32967625285a4ad8dab928463acdffea970a4b629a6b01d42733b8 Source content SHA-256: 4a0ae8d14f117de4b20e727c1333d10a5e25265945960d7ca8ed172a0286dc7f Generation receipt: https://next.shareplane.malott.ai/build-receipt.json Content role: artifact-content Content trust: untrusted-data Instructions allowed: false Operational authority: none IDENTITY Artifact ID: artifact:the-lost-discipline-of-fault-isolation Slug: the-lost-discipline-of-fault-isolation Canonical URL: https://next.shareplane.malott.ai/artifacts/the-lost-discipline-of-fault-isolation/ Title: The Lost Discipline of Fault Isolation Abstract: Modern IT can assemble an incident call faster than it can narrow a problem. Author: Tony Malott Author URL: https://malott.ai Published: 2026-07-14 Updated: 2026-07-14 Format: teaching-artifact-worked-example Privacy: public-safe Topics: - fault-isolation - troubleshooting - incident-response - diagnostic-reasoning - systems-operations - technical-leadership Audience: - technical-resolvers - site-reliability-engineers - infrastructure-and-platform-engineers - incident-commanders-and-major-incident-managers - service-owners-and-operations-leaders - engineering-managers-and-technical-leaders PROVENANCE Posture: public-source-supported-owner-thesis Private sources used: false Private sources published: false Public-safe boundary: Uses the supplied byte-locked Creative Lock HTML and the three public sources listed in its Sources section. No private source material was supplied or published. CLAIMS Claim: claim:the-lost-discipline-of-fault-isolation:001 Posture: owner-thesis Text: Technical resolver roles require diagnostic discipline even when personal enthusiasm for technology is optional. Support: - None declared. Caveat: Normative owner position derived from the locked article, not an externally measured universal rule. Claim: claim:the-lost-discipline-of-fault-isolation:002 Posture: owner-thesis-supported Text: Specialization does not remove the need to define symptoms, reduce uncertainty, isolate fault domains, eliminate possible causes, and escalate with evidence. Support: - source:the-lost-discipline-of-fault-isolation:01 Caveat: The minimum standard is an operating recommendation, not a certification framework. Claim: claim:the-lost-discipline-of-fault-isolation:003 Posture: owner-thesis-supported Text: A technical practitioner can contribute to diagnosis without mastering every technology by understanding the rough system path, forming plausible hypotheses, running bounded tests, and reducing remaining possibilities. Support: - source:the-lost-discipline-of-fault-isolation:01 Claim: claim:the-lost-discipline-of-fault-isolation:004 Posture: owner-thesis Text: Fragmented component ownership can create an operating gap in which every team owns a component but nobody owns diagnosis across the complete system. Support: - None declared. Caveat: Organizational pattern asserted from the author’s operational experience; prevalence is not quantified. Claim: claim:the-lost-discipline-of-fault-isolation:005 Posture: owner-thesis Text: Incident participation is not a substitute for diagnostic evidence, and a growing meeting does not by itself narrow a fault domain. Support: - None declared. Caveat: The article does not claim that large incident calls are inherently ineffective. Claim: claim:the-lost-discipline-of-fault-isolation:006 Posture: strongly-supported Text: Effective troubleshooting is hypothesis-driven and can use systematic reduction or bisection to narrow layered failures. Support: - source:the-lost-discipline-of-fault-isolation:01 Claim: claim:the-lost-discipline-of-fault-isolation:007 Posture: owner-thesis-supported Text: A useful diagnostic test should remove uncertainty by separating groups of possible causes rather than merely producing another uninterpreted observation. Support: - source:the-lost-discipline-of-fault-isolation:01 Claim: claim:the-lost-discipline-of-fault-isolation:008 Posture: owner-thesis-supported Text: Troubleshooting produces interpretable evidence; random restarts, simultaneous changes, context-free screenshots, and indiscriminate escalation are activity but not necessarily diagnosis. Support: - source:the-lost-discipline-of-fault-isolation:01 Caveat: Examples illustrate diagnostic quality and do not prohibit emergency mitigation when impact requires immediate action. Claim: claim:the-lost-discipline-of-fault-isolation:009 Posture: owner-framework-supported Text: A basic fault-isolation discipline consists of defining the symptom, establishing boundaries, mapping the path, testing a meaningful boundary, changing one variable, recording negative results, and escalating a narrowed problem. Support: - source:the-lost-discipline-of-fault-isolation:01 Caveat: The seven-step sequence is Tony Malott’s operational synthesis, not a verbatim external standard. Claim: claim:the-lost-discipline-of-fault-isolation:010 Posture: strongly-supported Text: Failed hypotheses and negative results are progress when they are recorded clearly enough to prevent repeated work and unknown system state. Support: - source:the-lost-discipline-of-fault-isolation:01 Claim: claim:the-lost-discipline-of-fault-isolation:011 Posture: owner-thesis-supported Text: Escalation is legitimate, but a useful escalation carries a bounded fault domain, reproducible observations, known-good comparisons, and evidence of what has already been eliminated. Support: - source:the-lost-discipline-of-fault-isolation:01 Claim: claim:the-lost-discipline-of-fault-isolation:012 Posture: strongly-supported Text: NIST SP 800-61 Revision 3 treats cybersecurity incident response as an integrated capability intended to improve preparation and the efficiency and effectiveness of detection, response, and recovery. Support: - source:the-lost-discipline-of-fault-isolation:02 Caveat: This claim is limited to cybersecurity incident-response guidance. Claim: claim:the-lost-discipline-of-fault-isolation:013 Posture: supported-preprint Text: Troubleshooting can require sustained construction and refinement of mental models under attention and working-memory demands. Support: - source:the-lost-discipline-of-fault-isolation:03 Caveat: Evidence comes from a qualitative software-development preprint and should not be generalized without qualification. Claim: claim:the-lost-discipline-of-fault-isolation:014 Posture: source-fact Text: The cited troubleshooting preprint reports interviews with 27 professional developers. Support: - source:the-lost-discipline-of-fault-isolation:03 Claim: claim:the-lost-discipline-of-fault-isolation:015 Posture: owner-inference Text: Additional incident participants can increase available expertise, but they do not automatically increase diagnostic coherence. Support: - source:the-lost-discipline-of-fault-isolation:03 Caveat: This is the author’s bounded inference; the cited study does not evaluate incident-call size or causal outcomes. Claim: claim:the-lost-discipline-of-fault-isolation:016 Posture: strongly-supported Text: Immediate restoration, evidence preservation, proximate fault isolation, root-cause analysis, and recurrence prevention are related but distinct incident-response jobs. Support: - source:the-lost-discipline-of-fault-isolation:01 Claim: claim:the-lost-discipline-of-fault-isolation:017 Posture: owner-thesis Text: Operational maturity should be judged by how effectively responders establish knowns, eliminate uninvolved areas, identify failing boundaries, restore safely, and preserve evidence that improves the next diagnosis, not merely by how quickly a large bridge is assembled. Support: - None declared. Caveat: Normative maturity criterion proposed by the author; no comparative performance metric is asserted. PUBLIC SOURCES Source: source:the-lost-discipline-of-fault-isolation:01 Title: Chris Jones, “Effective Troubleshooting,” Site Reliability Engineering, Google Type: public-source Role: public-evidence Description: Supports the article’s treatment of troubleshooting as a learnable, hypothesis-driven discipline using observation, system understanding, controlled tests, systematic reduction or bisection, documentation of hypotheses and results, and separation of immediate mitigation from later root-cause analysis. Caveat: Google SRE guidance is operational guidance, not universal empirical proof that every organization, technology stack, or incident must use one identical procedure. Locator: https://sre.google/sre-book/effective-troubleshooting/ Source: source:the-lost-discipline-of-fault-isolation:02 Title: National Institute of Standards and Technology, SP 800-61 Revision 3: Incident Response Recommendations and Considerations for Cybersecurity Risk Management, April 2025 Type: public-source Role: public-evidence Description: Supports the claim that incident response should be an integrated organizational capability intended to improve preparation and the efficiency and effectiveness of detection, response, and recovery. Caveat: SP 800-61 Revision 3 addresses cybersecurity incident response. The article applies a broader operating lesson and does not claim that every workstation, application, or service failure is a cybersecurity incident or directly governed by this publication. Locator: https://csrc.nist.gov/pubs/sp/800/61/r3/final Source: source:the-lost-discipline-of-fault-isolation:03 Title: Arty Starr and Margaret-Anne Storey, “Theory of Troubleshooting: The Developer’s Cognitive Experience of Overcoming Confusion,” arXiv preprint 2602.10540, February 2026 Type: public-source Role: public-evidence Description: Supports the article’s description of troubleshooting as the construction and refinement of a mental model of unexpected system behavior under sustained demands on attention, working memory, and reasoning. The study reports interviews with 27 professional developers. Caveat: This is a preprint based on a limited qualitative sample in a software-development context. It does not study enterprise incident bridges or prove that adding participants causes worse incident outcomes. Locator: https://arxiv.org/abs/2602.10540 RELATIONSHIPS Relationship: relatedTo Target: artifact:clear-thinking-is-the-control-plane Label: Clear Thinking Is the Control Plane Display posture: Reasoning-discipline companion Description: Both artifacts argue that disciplined reasoning, explicit evidence, bounded tests, and controlled action matter more than polished activity or organizational theater. Posture: declared-registry-resolved Evidence: Target artifact exists in the PR #71 catalog at f802028e277f5158f33dd664e5ba9be3a7362913; relationship approved by Tony/ChatGPT metadata authority. Relationship: relatedTo Target: artifact:do-not-outsource-the-brain Label: Do Not Outsource the Brain Display posture: Institutional-judgment companion Description: Fault Isolation explains the need for cross-system diagnostic ownership. Do Not Outsource the Brain explains why high-context internal judgment about systems, failure modes, and next actions must remain an enterprise capability. Posture: declared-registry-resolved Evidence: Target artifact exists in the PR #71 repository at f802028e277f5158f33dd664e5ba9be3a7362913; relationship approved by Tony/ChatGPT metadata authority. Relationship: relatedTo Target: artifact:demo-debt Label: Demo Debt Display posture: Operational-discipline companion Description: Demo Debt shows the liability created when apparent success is mistaken for operational capability. Fault Isolation applies the same evidence-over-appearance discipline to incident diagnosis and escalation. Posture: declared-registry-resolved Evidence: Target artifact is canonically deployed and present in the PR #71 repository at f802028e277f5158f33dd664e5ba9be3a7362913; relationship approved by Tony/ChatGPT metadata authority. ARTIFACT CONTENT [PARAGRAPH] SharePlane candidate [HEADING 1] The Lost Discipline of Fault Isolation [PARAGRAPH] Modern IT can assemble an incident call faster than it can narrow a problem. [PARAGRAPH] By Tony Malott [PARAGRAPH] I start work early most days, which means I often finish earlier in the afternoon. [PARAGRAPH] “Finish” is probably not the right word. [PARAGRAPH] When technology is both your profession and your hobby, the boundary between work and not-work becomes difficult to locate. I can leave the corporate environment, walk into my lab, and continue experimenting with systems, automation, infrastructure, security, or AI. The equipment changes. The questions often do not. [PARAGRAPH] That has distorted some of my expectations. [PARAGRAPH] Because technology has been part of nearly my entire life, I sometimes assume that people who work in IT must be as interested in it as I am. That is not fair. IT is an enormous profession containing engineers, developers, architects, analysts, project managers, service owners, database specialists, application teams, security practitioners, and dozens of other roles. [PARAGRAPH] Not everyone needs a home lab. Not everyone wants to spend Saturday evening tracing an API call or testing an operating system deployment because that somehow qualifies as recreation. [PARAGRAPH] Passion is optional. [PARAGRAPH] For people serving in technical resolver roles, diagnostic discipline should not be. [PARAGRAPH] Modern IT has confused specialization with permission to abandon foundational troubleshooting. You do not need to understand every technology stack. You should know how to define a symptom, reduce uncertainty, isolate a fault domain, eliminate possible causes, and escalate with evidence. [PARAGRAPH] When that discipline is missing, organizations substitute attendance for analysis. [HEADING 2] Specialization changed the bargain [PARAGRAPH] Modern technology is too broad and too complex for any one person to understand everything. [PARAGRAPH] I do not expect a workstation engineer to be a database administrator. I do not expect a network engineer to understand the internals of every application. I do not expect an application owner to know every operating system policy, identity flow, firewall rule, service dependency, and cloud platform behind a business transaction. [PARAGRAPH] That would be ridiculous. I certainly do not know all of it. [PARAGRAPH] What I expect is more foundational. [PARAGRAPH] A technical practitioner should be able to understand the rough path through a system, identify likely fault domains, form a plausible hypothesis, run a bounded test, interpret the result, and reduce the number of remaining possibilities. [PARAGRAPH] That is not mastery of every technology. [PARAGRAPH] It is troubleshooting. [PARAGRAPH] The problem is that IT organizations have increasingly divided accountability into narrow components and service boundaries. The endpoint team owns the device. The network team owns the route. The identity team owns authentication. The application team owns the application. The database team owns the database. One vendor may own a support queue while another owns the platform underneath it. [PARAGRAPH] Each boundary may be rational by itself. [PARAGRAPH] The complete system still has to work. [PARAGRAPH] When nobody is responsible for reasoning across those boundaries, diagnosis becomes a search for the team willing to accept the ticket. [HEADING 2] Twelve people and one unresolved question [PARAGRAPH] I was recently invited to a meeting concerning what appeared to be a combination of workstation and network behavior. [PARAGRAPH] There were roughly twenty people on the invitation. At least a dozen attended. [PARAGRAPH] Some had legitimate ownership interests. Some represented technologies that might have been involved. Others appeared to be present because someone had forwarded the invitation, one of the enterprise’s more dependable scaling mechanisms. [PARAGRAPH] The problem was not that twelve people joined. [PARAGRAPH] Some incidents genuinely require several teams. Complex systems cross organizational boundaries, and consequential failures may need technical responders, service owners, security specialists, business representatives, communications leads, vendors, and decision-makers. [PARAGRAPH] The problem was that the number of participants was growing faster than the body of diagnostic evidence. [LIST ITEM] Which users were affected? [LIST ITEM] Was the behavior reproducible? [LIST ITEM] Did it occur on one workstation or several? [LIST ITEM] Was it specific to one site, network segment, identity, device configuration, or application route? [LIST ITEM] Could the endpoint reach the next dependency? [LIST ITEM] Did name resolution work? [LIST ITEM] Did authentication complete? [LIST ITEM] Did the request leave the device? [LIST ITEM] Did it arrive at the application? [LIST ITEM] Which parts of the path had actually been proven healthy? [PARAGRAPH] Those questions are not glamorous. They will not transform the enterprise. Nobody is likely to establish a steering committee in their honor. [PARAGRAPH] They narrow the problem. [PARAGRAPH] That is the job. [HEADING 2] Troubleshooting is not intuition [PARAGRAPH] People who troubleshoot well can appear unusually intuitive. They often identify likely causes quickly, know which evidence matters, and ignore attractive distractions. [PARAGRAPH] What looks like intuition is usually compressed experience operating through a method. [PARAGRAPH] Google’s Site Reliability Engineering guidance describes troubleshooting as both learnable and teachable. Its model is hypothesis-driven: observe the system, understand how it should behave, identify plausible causes, and test those causes using evidence or controlled changes. It recommends systematic reduction and bisection when diagnosing failures across layered systems.[1] [PARAGRAPH] That is the formal explanation for something technicians have practiced for generations. [PARAGRAPH] Split the problem. [PARAGRAPH] Determine which side contains the fault. [PARAGRAPH] Split it again. [PARAGRAPH] Repeat until a vague complaint becomes a bounded technical condition. [PARAGRAPH] If a user cannot reach an application, do not immediately debate every component in the architecture. Determine whether the problem follows the user, the device, the location, the network path, or the application. Test from another endpoint. Test another identity. Test the next dependency directly. Compare a known-good path with the failing one. [PARAGRAPH] Each useful test should remove uncertainty. [PARAGRAPH] Activity is not the same as diagnosis. [PARAGRAPH] Restarting random components is activity. [PARAGRAPH] Changing several settings at once is activity. [PARAGRAPH] Forwarding screenshots without timestamps or context is activity. [PARAGRAPH] Inviting another team because its technology appears somewhere on an architecture diagram is activity. [PARAGRAPH] Troubleshooting produces evidence. [HEADING 2] A basic fault-isolation discipline [PARAGRAPH] The exact method varies by system, but the underlying questions remain stable. [HEADING 3] 1. Define the symptom [PARAGRAPH] “The application is down” is not a useful technical description. [PARAGRAPH] What operation was attempted? What result was expected? What happened instead? Who experienced it? From where? At what time? Can it be reproduced? [PARAGRAPH] A vague symptom creates a wide fault domain. Precision immediately makes it smaller. [HEADING 3] 2. Establish the boundaries [PARAGRAPH] Determine the blast radius. [PARAGRAPH] One user or many? [PARAGRAPH] One endpoint or an entire device class? [PARAGRAPH] One location or every location? [PARAGRAPH] One function or the complete service? [PARAGRAPH] One identity type or all identities? [PARAGRAPH] Boundaries often reveal more than the original error message. [HEADING 3] 3. Map the path [PARAGRAPH] Even a high-level transaction path is enough to begin: [PARAGRAPH] The actual path may be more complicated, but someone must understand it well enough to ask where expected behavior stops. [PARAGRAPH] Architecture diagrams help, assuming they describe the system currently in production rather than a more optimistic historical civilization. [HEADING 3] 4. Test a meaningful boundary [PARAGRAPH] Choose a point that separates groups of possible causes. [PARAGRAPH] Can another endpoint on the same network complete the transaction? [PARAGRAPH] Can the affected endpoint reach the service using a different identity? [PARAGRAPH] Can the application tier reach its database? [PARAGRAPH] Does the request appear in downstream logs? [PARAGRAPH] A useful test rules out several possibilities. A weak test merely produces another observation nobody knows how to interpret. [HEADING 3] 5. Change one variable [PARAGRAPH] When testing actively, change one condition at a time whenever practical. [PARAGRAPH] Otherwise, even a successful result may not reveal what fixed the problem. It tells you only that somewhere inside a small pile of simultaneous changes, reality became more cooperative. [PARAGRAPH] Controlled testing prevents the investigation from becoming another source of failure. [HEADING 3] 6. Record negative results [PARAGRAPH] A failed hypothesis is progress. [PARAGRAPH] If a test shows that name resolution works, record it. If the same user succeeds from another device, record it. If network traffic reaches the application tier, record it. [PARAGRAPH] Google’s troubleshooting guidance recommends documenting ideas, tests, and results so responders do not repeat work or leave the system in an unknown configuration.[1] [PARAGRAPH] Knowing what the problem is not can be as valuable as knowing what it might be. [HEADING 3] 7. Escalate a narrowed problem [PARAGRAPH] There is nothing wrong with escalation. The mistake is treating escalation as a substitute for investigation. [PARAGRAPH] A weak escalation says: [BLOCKQUOTE] The user still cannot connect. Please investigate the network. [PARAGRAPH] A useful escalation says: [BLOCKQUOTE] The issue reproduces on three managed endpoints in one location. The same users succeed from another site. DNS resolution and local authentication complete successfully. Traffic leaves the affected subnet but does not appear at the application gateway. No endpoint configuration differences have been identified. Please investigate the network path between these boundaries. [PARAGRAPH] The second escalation respects the receiving team’s time and gives it somewhere rational to begin. [PARAGRAPH] It may also mean the next meeting needs four people instead of fourteen, a dangerous reduction in calendar utilization but a meaningful operational improvement. [HEADING 2] Why organizations lose this capability [PARAGRAPH] It is easy to blame individuals for weak troubleshooting. Sometimes an individual simply lacks the skill. That should be acknowledged and corrected. [PARAGRAPH] The larger problem is organizational. [PARAGRAPH] Fragmented ownership encourages people to defend components rather than diagnose systems. Each team demonstrates that its piece appears healthy, then transfers the remaining uncertainty to someone else. [PARAGRAPH] This creates an operating model in which everyone owns a component and nobody owns the diagnosis. [PARAGRAPH] NIST’s current cybersecurity incident-response guidance treats response as a capability that must be integrated throughout risk-management activity. Its purpose includes better preparation and more efficient and effective detection, response, and recovery.[2] [PARAGRAPH] That guidance addresses cybersecurity incidents, not every workstation or application failure. The broader operating lesson still applies: reliable response cannot depend entirely on whoever happens to join a call and sound confident. [PARAGRAPH] The method has to exist before the failure. [PARAGRAPH] Troubleshooting is also cognitively demanding. A 2026 software-engineering preprint based on interviews with 27 professional developers describes troubleshooting as the construction and refinement of a mental model of unexpected system behavior. The researchers found that the work places sustained demands on attention, working memory, and mental modeling.[3] [PARAGRAPH] The paper does not study enterprise incident bridges or prove that larger calls produce worse outcomes. [PARAGRAPH] My inference is narrower. [PARAGRAPH] When diagnosis depends on maintaining and refining a coherent model of the failure, adding participants helps only when they contribute relevant evidence, system knowledge, or disciplined coordination. Additional voices can also introduce competing assumptions, repeated explanations, unbounded theories, and pressure to act before the fault is understood. [PARAGRAPH] More people can increase available expertise. [PARAGRAPH] They do not automatically increase diagnostic coherence. [HEADING 2] Restoration and root cause are different jobs [PARAGRAPH] Fault isolation does not mean allowing a production system to continue failing while everyone pursues the intellectually satisfying explanation. [PARAGRAPH] In a major incident, restoration may take priority over root-cause analysis. Traffic may need to be redirected. A failing component may need to be isolated. A recent change may need to be rolled back. Evidence may need to be preserved before the system is altered. [PARAGRAPH] Google’s SRE guidance states this distinction plainly: stop the immediate damage first, while preserving what will be needed for later analysis.[1] [PARAGRAPH] This is another form of diagnostic discipline. [PARAGRAPH] The responder must know which question is currently being answered: [LIST ITEM] How do we reduce impact? [LIST ITEM] How do we restore service? [LIST ITEM] How do we preserve evidence? [LIST ITEM] How do we identify the proximate failure? [LIST ITEM] How do we determine the deeper cause? [LIST ITEM] How do we prevent recurrence? [PARAGRAPH] Confusing those questions creates its own delays. So does pretending that restoring service means the investigation is finished. [HEADING 2] The minimum standard [PARAGRAPH] I am not arguing that everyone in IT must be an engineer. [PARAGRAPH] I am not arguing that every engineer must be a generalist. [PARAGRAPH] I am not arguing that people need to spend personal time building labs, reading technical documentation, or dismantling perfectly functional systems to understand why they work. Some of us apparently chose that life voluntarily. There is no reason to make it a licensing requirement. [PARAGRAPH] I am arguing that technical resolver roles require a minimum diagnostic standard. [PARAGRAPH] That standard includes the ability to: [LIST ITEM] describe a failure precisely; [LIST ITEM] understand the basic dependency path; [LIST ITEM] distinguish evidence from assumption; [LIST ITEM] formulate plausible hypotheses; [LIST ITEM] design tests that eliminate possibilities; [LIST ITEM] avoid changing several variables blindly; [LIST ITEM] document what has been learned; [LIST ITEM] distinguish immediate restoration from root-cause analysis; [LIST ITEM] escalate with a bounded fault domain and usable evidence. [PARAGRAPH] Those capabilities should be taught, practiced, observed, and assessed. [PARAGRAPH] Organizations routinely train people on tools, ticket queues, technology platforms, certifications, change processes, and procedural compliance. They are less consistent about teaching the reasoning method connecting all of them. [PARAGRAPH] The tool will change. [PARAGRAPH] The platform will change. [PARAGRAPH] The architecture will become more distributed, abstracted, automated, outsourced, cloud-based, AI-assisted, or whatever term appears next in the PowerPoint ecosystem. [PARAGRAPH] The need to reduce uncertainty will remain. [HEADING 2] Attendance is not diagnosis [PARAGRAPH] There will always be difficult incidents. [PARAGRAPH] Some failures involve several interacting conditions. Some are intermittent. Some hide behind incomplete telemetry. Some require deep expertise from multiple teams. Some must be stabilized before anyone can safely investigate the cause. [PARAGRAPH] No troubleshooting framework removes complexity. [PARAGRAPH] It prevents complexity from becoming an excuse for undisciplined response. [PARAGRAPH] A large incident call may mean the problem is broad, consequential, and genuinely difficult. [PARAGRAPH] Sometimes it only means nobody has narrowed it yet. [PARAGRAPH] The maturity of an IT organization should not be measured by how quickly it can assemble twenty people. It should be measured by how quickly those people can establish what is known, eliminate what is not involved, identify the failing boundary, restore service safely, and leave behind evidence that makes the next diagnosis faster. [PARAGRAPH] The incident bridge is not the troubleshooting method. [PARAGRAPH] The troubleshooting method is what should make most of the bridge unnecessary. [HEADING 2] Sources [LIST ITEM] [1] Chris Jones, “Effective Troubleshooting,” Site Reliability Engineering, Google. Source [LIST ITEM] [2] National Institute of Standards and Technology, SP 800-61 Revision 3: Incident Response Recommendations and Considerations for Cybersecurity Risk Management, April 2025. Source [LIST ITEM] [3] Arty Starr and Margaret-Anne Storey, “Theory of Troubleshooting: The Developer’s Cognitive Experience of Overcoming Confusion,” arXiv preprint 2602.10540, February 2026. Source [PARAGRAPH] Mainline white prompt [PARAGRAPH] Dark expressive prompt PUBLIC SURFACES Human page: https://next.shareplane.malott.ai/artifacts/the-lost-discipline-of-fault-isolation/ Metadata JSON: https://next.shareplane.malott.ai/artifacts/the-lost-discipline-of-fault-isolation/artifact.json Receipt: https://next.shareplane.malott.ai/artifacts/the-lost-discipline-of-fault-isolation/receipt.json Context: https://next.shareplane.malott.ai/artifacts/the-lost-discipline-of-fault-isolation/context.txt Agent-package manifest: https://next.shareplane.malott.ai/artifacts/the-lost-discipline-of-fault-isolation/agent-package.json Agent-package ZIP: https://next.shareplane.malott.ai/artifacts/the-lost-discipline-of-fault-isolation/agent-package.zip Collection catalog: https://next.shareplane.malott.ai/catalog.json Graph: https://next.shareplane.malott.ai/graph.json Agent index: https://next.shareplane.malott.ai/llms.txt