Advice to a new CIO
Observe, Understand, Govern, Adapt
So many of the pieces I post here speak to security and privacy professionals, offering advice on how to manage ‘up’. But often when working with CIOs I find the questions they have equally challenging. Whether you’re new to the CIO role, or at a new institution, starting a position is a terrific opportunity. I’ve tried to use every position change as both an opportunity to reinvent myself (to reflect on my own missteps and shortcomings and try to avoid them going forward) as well as to reframe for my staff and colleagues my approach to infosec and being an institutional actor in general.
For a new CIO to whom infosec probably reports - at least in higher ed - there is also a window of opportunity to have a strategic impact on your infosec program. Now, I can just hear everyone saying “you dope, shouldn’t you just listen to your CISO et al and figure out the lay of the land before requesting changes in the zoning laws?” Of course, but the questions you ask about the infosec program, those very first questions on day one, truly have a long-tail of impact. They’re a tell, your new staff will read and probably repeat to one another for years.
What I want to do is separate what you need to know as a manager with operational obligations from what’s valuable as an executive with strategic responsibilities. The former is going to be concerned with issues of execution: what are we doing, where are the compliance gaps, and what deployment constraints do we have? If your team has summarized these in their Information Assurance Management Plan or IAMP, great! You can read that over lunch and follow up with your CISO during regular meetings. If they haven’t created an IAMP or its equivalent, asking for one is a great step - though I’d do so after working through the more strategic questions.
It is these strategic questions that, as a campus executive, I would open with. As you might imagine, some deal with governance and decision making, but others try to get at many of the concerns I’ve raised throughout this blog. I’ve grouped the questions into four phases:
What is known? (Phase 1)
What is testable? (Phase 2)
What is enforceable? (Phase 3)
What is scalable? (Phase 4)
The bottom line is that while it’s easy to ask questions that the CISO can discuss but I’d like to turn them into something they must answer, with evidence.
Phase 1: The Visibility Diagnostic
In phase 1 we want to test whether the program is grounded in telemetry and outcomes, or in asserted capability derived from tool ownership. Essentially we’re going to ask a deceptively simple question: how do you know what you know? More specifically, is the security program grounded in observed system behavior, or in asserted capability derived from tool ownership? It’s far too easy (lord knows I’ve consistently played this card myself) when asked about, for example, email hygiene, to say “well, we have paid some absurd amount for Microsoft’s anti-phishing enhancements.” This is responding to a question on performance metrics with a tool inventory. So my suggested question would be, “what aspects of our security posture are directly measured through system behavior - and what are we inferring based solely on the controls we believe are in place?” This is a difficult question to answer unless performance outcomes are predefined, and answering it forces your CISO to split their answer into one of two buckets: measured vs. assumed. Hopefully it will create some immediate discomfort if everything falls into the second category. Most organizations discover, often uncomfortably, that a much larger portion of their confidence rests on assumption than they expected.
Along similar lines, you’ll want to push your team to define success in operational terms and surface performance gaps, not just activity metrics. With this next question you’ll see that I’ve added in the notion of “under stress.” This hearkens back to my concept of a behavioral architecture by forcing an admission of failure and not just a definition of success. “For our most critical controls, what explicit outcomes are they expected to achieve under stress - and where do we have evidence they are not meeting those expectations?” Notice the subtle shift. We're no longer asking whether we own a firewall or an EDR platform. We're asking what claims we are willing to make about their behavior when they are actually needed. Every engineering discipline eventually reaches this point. Bridges are specified by the loads they can carry, not simply by the fact that they contain steel. Aircraft are certified by their behavior under failure, not by an inventory of their components. Cybersecurity, I think, needs to become more comfortable speaking in that same language
Phase 2: The Architectural Diagnostic
I would note that one dimension to being a CISO is dealing with the multitude of pressures you face with regard to specific controls. Compliance regulations and organizational policy often impose controls and activities of low value for which the security organization has no choice but to implement. In some cases these might even damage your security posture - I’m thinking about obsolete password formation and change policies that auditors love to persist in embracing. Be that as it may, this reality makes it important to distinguish between a security program built on engineering and one built on belief. Every security team develops convictions about what “works.” The question is whether those convictions are actually testable.
The question I would ask to start examining this is, “which core assumptions about our security posture could be wrong today - and what signal would tell us they’ve already failed?” Now there’s a risk to this question. By asking “could be wrong today” and “already failed” there’s an immediacy to the question. Being told by your CISO that some security service, control, or function has already failed is a bell being rung you can’t unhear. It gets at the difficult reality that those expensive security solutions you fought to fund may in fact be obsolete. The treadmill keeps turning.
I have two more questions for this second phase, one on dependency mapping and structural blind spots, and another on feedback loops.
For the first, “can we trace a critical institutional service end-to-end to the infrastructure and controls it depends on - and identify the specific points where failure would interrupt it?”; the second asks, “what is a recent incident or near-miss that caused a permanent change to our architecture or control behavior - and what exactly changed?”
The dependency question tests whether the organization actually understands the systems it operates or merely the technologies it owns. Every mature environment contains hidden dependencies that only reveal themselves under stress. Naturally this works for interrogating your organization outside of security as well. It’ll help surface assumptions (such as once when a power outage took out our data center. The network folks were very proud that all the major nodes had backup power. Unfortunately DNS sat entirely in the primary DC which was down for eight hours. A fuller dependency map would have surfaced this.)
The second question is, I think, a bit more interesting. It determines whether incidents produce structural adaptation or only procedural updates. By asking, “what exactly changed,” you require a kind of precision that eliminates hand waving. It’s very common to look at incidents (be they malware, attacks, or simple IT failures) as point problems. What you’re hoping to engender with this question is for failures to result in a holistic understanding of your ecosystem. By insisting on “what exactly changed” the conversation shifts from documenting failure to engineering resilience.
Obviously I’m dancing around the behavioral architecture idea I introduced in my last post. Behavioral architectures are valuable precisely because they define expected behavior under failure. They don't simply enumerate controls; they specify how the system is expected to respond when assumptions break. Without explicit failure conditions, there is no meaningful architecture - only aspiration.
Phase 3: The Governance Diagnostic
Here I’m going to talk about risk decisions, but as I edit this I’m thinking the better way to frame this is around organizational behavior. Can this organization reliably convert decisions into operational change? I’ve long been struck by our obsession with risk decisions. Canonical thinking about information security rightfully pushes many critical risk decisions high up the organization’s orgchart, and for good reason: risk decisions always involve two elements: resources for mitigation and risk acceptance. Two elements that usually require very senior executives. I recall irritating a CIO I worked for when I insisted on taking an issue to the Provost since the resultant risk impacted the entire academic mission. I felt it was of such broad impact it went beyond the CIO’s scope.
Yet, as every CISO knows, we make dozens of risk and risk acceptance decisions daily. Hopefully we do so with enough savvy to recognize when we’re exceeding our own scope, but this is the reality of being in a security leadership role. Nevertheless, it’s valuable to link decisions and outcomes. So for phase 3 my opening question would be, “where are risk decisions actually made—and what is a recent example where that decision produced a measurable change in the environment?” Asking for an example moves this from theory to practice. Notice the emphasis on “actually”. Every organization has governance diagrams. Those diagrams tell you where authority is supposed to reside. Asking for a concrete example shifts the discussion from organizational charts to organizational behavior. It reveals whether governance is producing operational consequences or merely documenting conversations. Importantly if answered honestly it separates formal governance structure from actual decision authority and enforcement. It opens the door to that honest conversation about risk decisions and risk acceptance that separates the theatrical from the operational.
There’s a lot to plumb around the issue of risk acceptance - but I want to turn to what might be called “exception discipline” for the next question. “How do we track and review policy exceptions - and when was the last time we revoked or reversed one?” Fortunately, I’ve noticed more and more schools codifying a rather rigorous policy exception process; if you don’t have one, you’re behind the curve. For this question, the term “revoked” is the real test, most organizations can’t answer that. But the goal here is obviously one of maturation - you want to move from passive tracking to active discipline. This question tests whether exception management is active governance or passive accumulation.
I want to close out phase 3 by looking at the question of decentralization. “In our most decentralized environments, who is accountable for security outcomes - and how do we verify that accountability is being met?” Notice that I’m still pushing on outcomes over roles, and by using “verify” explicitly, we surface gaps between nominal ownership and enforceable accountability in a distributed environment. This seems like a fairly straightforward question, but in higher education it truly is a can of worms. While it focuses on accountability, it requires deep engagement on issues from outcome and role definitions, system telemetry and visibility, to resource allocation. I suspect anyone wrestling with the management of risk in distributed environments (or defining the distributed IT role in research security) could do much worse than starting here.
Remember, exceptions are organizational entropy. Every exception is locally rational. Yet it’s tempting to say that collectively they redefine the architecture.
Phase 4: The Strategic Diagnostic
The previous phases examined whether the security program understands itself today. The final phase asks a different question: will it still work tomorrow? Strategy is fundamentally about preparing for conditions that have not yet arrived. That means understanding not only where the program succeeds today, but where it will fail as demand, regulation, and adversary capability continue to accelerate.
As regular readers of this blog know, I’m always going to advocate for collective action. When I ask CISOs (or security practitioners in general) about what prevents them from collaborating, the most common answer is that they don’t feel encouraged to do so. What they hear from their management is, “why should we invest your time in helping others when our own challenges are so vast?” is commonly reported. I do think some of this is self-imposed. For the most part, collaboration beyond commiseration isn’t well supported in our community. We seem to lack much of the infrastructure for true collaboration. That makes it difficult to give “just a bit of time.” It’s the self-fulfilling prophecy of not having the time to collectively solve problems so we spend even more time solving them independently. Which is why I think this next question is so essential if we’re forward-looking: “which of our major security challenges are structurally shared with our peers - and what are we doing today to solve them collectively rather than locally?” This tests whether the organization is leveraging a shared problem space or defaulting to isolated solutions. And no, being told “I attended a talk on how some other school tackled this” is not a great response. I recognize that it’s very difficult for a smaller, less resourced school to participate in collaborative projects. The temptation of simply identifying a successful peer and emulating them is overwhelming, and frankly, probably the right move. But even in those circumstances, it is essential that those smaller schools to participate, however modestly, if for no other reason than to make sure their needs are included as requirements. Collaboration by only the wealthy is a country club, not a model for higher education.
Finally, I want to return to the issue of system failures. “Under a step-function increase in demand or threat activity, where does our current model fail first - and how do we know?” I like this question because it forces several admissions. First, the team must identify actual breaking points rather than vague concerns. Every system has a limiting reagent - whether it’s staffing, identity infrastructure, governance throughput, incident response capacity, or simply budget. Second, it requires the organization to articulate what it believes its operating model actually is. Surprisingly few security programs possess an explicit model. Most possess diagrams.
With that, the real opportunity arises. Does that model just show what tools you have, i.e., components and capabilities? Does it simply show where those tools and controls operate, i.e., the topology of your security model? What you’ll want to see is some system-oriented, behavioral thinking. Identifying indicators leading to outcomes: from what exists to what must happen.
I can think of a number of alternative ways to structure the questions, and with it, the arc of your organizational diagnostic. But I like the four listed above, perhaps slightly reworded as,
Observe (telemetry) - can you observe it?
Understand (architecture) - can you explain it?
Govern (decision-making) - can you act on it?
Adapt (strategy) - will it continue to work when the environment changes?
Which feels to me like more than merely a diagnostic, but a repeatable model.
Ultimately, these questions aren't intended to evaluate a CISO's technical knowledge or even their management of the security function. They're intended to reveal whether the security program has crossed the line from managing technologies to engineering organizational behavior. Every phase pushes in the same direction: away from inventories and toward outcomes; away from asserted capability and toward observable behavior; away from static architecture and toward systems that continue to function under stress. That's the distinction between owning security products and operating a security program.


