Security Operations
Updated 11 min read

SOC Hiring, Onboarding, and Tier Structure That Actually Works

47%
Annual analyst attrition at poorly tuned SOCs
90
Days to full independent triage
3
Tiers in a typical mature SOC
20%
L2 time on detection engineering rotation

SponsoredRetool

Retool's new app builder is where AI-generated code ships safely

Building apps with AI is easy. Getting them to production safely is another story.

Start building for free today

Most SOC org charts are inherited rather than designed. A team adopts a three-tier model because the framework documents show three tiers, hires for tool certifications because the job board templates list them, and onboards by handing a new analyst a console login and a stack of runbooks. Six months later the senior leader is wondering why MTTD is stagnant, why three analysts left in a quarter, and why the L1 queue keeps growing.

This guide is for the leader rebuilding a SOC or scaling one past the breaking point. We cover tier model design with explicit headcount triggers, hiring criteria that predict performance rather than fit a template, a 90-day onboarding program built around graduated independence, the metrics that actually evaluate analyst contribution, and the structural changes that reduce attrition. The opinions here are strong because the alternative, generic advice, has produced the SOC fatigue crisis we are now in.

A modern SOC is not a triage factory. It is a detection capability with a triage function attached. Designing it the other way around is the most common and most consequential mistake leaders make.

Tier Model Design and When to Add Each Layer

Start with L1 as a triage function: alert acknowledgment, runbook execution, basic enrichment, and escalation to L2 with a populated investigation template. L1 work is bounded by documented procedures; if an L1 is making judgment calls outside the runbook, your runbook is incomplete or the alert should not be at L1. L2 is investigation: pivot across data sources, correlate, perform containment actions within defined authority, and produce an incident write-up. L2 needs context that L1 should not, including write access to response tooling and the authority to isolate hosts. L3 is detection engineering and threat hunting: no queue, no scheduled triage, focus on proactive work, rule development, tuning, threat modeling, and complex investigations escalated by L2. The headcount triggers are roughly these. Add L2 when L1 is escalating more than 20 percent of alerts and senior analysts are doing triage to keep the queue clean; this is the signal that you need an investigation layer. Add L3 when L2 is consistently above 80 percent utilization and detection content is going stale because nobody owns it. Below about 12 analysts you usually cannot afford true L3 separation; instead, dedicate 20 percent of senior L2 time to detection engineering on a rotation. Above 20 analysts the L3 function should be a distinct team with its own leader, because the work and incentives are different from triage.

Hiring Criteria That Predict Performance

The standard SOC job description optimizes for the wrong signals. Certifications correlate weakly with on-the-job performance after the first 90 days; they tell you a candidate can study, not that they can investigate. Tool-specific experience is overweighted; a strong analyst learns a new console in a week and the time horizon of a hire is years. What predicts performance: demonstrated investigation skill, writing ability, and intellectual curiosity. Test these directly. Give every candidate a real alert in the interview, with realistic context and ambiguity, and ask them to walk through their triage. You are watching for hypothesis generation, pivot selection, and willingness to say they do not know rather than confabulate. Have them write a short incident summary; clarity of writing is one of the strongest predictors of L2 effectiveness and almost nobody screens for it. Ask what they have learned outside work in the last six months and how; the answer reveals whether the curiosity is sustainable or performative. Drop the certification gate for entry-level roles, especially for candidates from adjacent fields such as system administration, network engineering, or technical support; these populations consistently outperform straight-from-bootcamp hires after the first year. Diversify your sourcing pools accordingly.

Free daily briefing

Briefings like this, every morning before 9am.

Threat intel, active CVEs, and campaign alerts, distilled for practitioners. 50,000+ subscribers. No noise.

The 90-Day Onboarding Program

Onboarding is where most SOCs fail their new hires, and the cost is invisible because the analyst stays but never reaches full effectiveness. Structure a 90-day program with graduated independence. Week one is shadow only: the new analyst sits with senior analysts during triage, asks questions, reads investigations, and writes nothing of consequence. The goal is mental model formation. Weeks two through four are supervised triage: the analyst works alerts but every action is reviewed by a buddy before submission. The buddy is a peer L2, not the manager, because the feedback loop needs to be high-frequency and low-stakes. Month two is independent triage with L2 review on a sample, say 20 percent of closed alerts, plus all escalations. Month three is full independent operation with metrics tracking and a 30-day retrospective. Require the new analyst to produce documentation continuously: every alert type they triage gets a runbook contribution, every gap they hit gets an open question logged. This is a forcing function for knowledge capture and surfaces gaps in your existing documentation faster than any audit. Assign a named onboarding owner who is not the manager; this protects the candor of the feedback loop and gives a senior analyst a development opportunity in mentorship.

Metrics That Evaluate Analyst Contribution Honestly

Tickets closed per day is the most common SOC metric and the most actively harmful. It rewards velocity over judgment and creates a structural incentive to close ambiguous alerts as false positive. Replace it with a small set of metrics that actually measure value. Mean time to detect and mean time to respond are team-level operational metrics; they should be tracked but not used to evaluate individuals because they are influenced by alert mix and shift composition. At the analyst level, look at false positive rate per analyst as a proxy for tuning skill; analysts who never propose a tuning change but close many alerts as false positive are recreating the same triage work daily and adding no leverage. Track escalation rate to detect both under-escalation, the analyst who closes things they should escalate, and over-escalation, the analyst who lacks confidence. Measure detection coverage contributed: rules written or improved, runbooks added, threat hunting reports produced. For L2 and above, this is the primary career metric. Review metrics quarterly with each analyst and use them as a conversation starter, not a performance trigger; mechanical management by metric is how you accelerate attrition.

Reducing Analyst Attrition Structurally

Analyst attrition above 30 percent annually is almost always a structural problem, not a hiring problem. The top driver across mature SOCs is alert fatigue from poorly tuned detections. If your analysts are working 80 percent false positive rates, no compensation package will retain them past 18 months. Fix the detection content. Track false positive rate per rule, retire or tune any rule above 90 percent false positive within 30 days of identification, and treat tuning as production work with its own queue. The second driver is lack of growth path. An L1 who sees no route to L2 and beyond will leave for the route at a different employer. Document the path explicitly: skills required, exposure required, evaluation cadence. Build a detection engineering rotation that brings L2 analysts into rule development for one week per month or 20 percent time; this provides both immediate retention value and a feeder pipeline into L3 when the role opens. The third driver is shift structure for follow-the-sun coverage; rotating night shifts are a known burnout accelerator and should be replaced where possible with regional follow-the-sun staffing or with dedicated night-shift hires compensated accordingly. Compensation matters, but it is rarely the primary lever; structural fixes outperform raises for retention.

MDR Integration: When to Outsource and How

Outsourcing tier one to an MDR provider is a legitimate strategy for teams that cannot staff 24x7 coverage in-house or cannot afford the management overhead of a global SOC. It is a poor strategy if used to avoid building security operations capability entirely. The right model is selective outsourcing: the MDR handles initial triage and well-defined response actions during off-hours and overflow, while your internal team retains detection engineering, threat hunting, complex investigations, and customer-specific context. Structure the contract around outcomes, not activity: time-to-acknowledge, time-to-escalate-with-context, percent of escalations with complete investigation packages, and rule tuning recommendations contributed per quarter. Avoid contracts that measure alerts handled, which incentivizes volume over judgment. Define the escalation criteria explicitly and in writing: what the MDR can close, what they must escalate, what they can contain without approval. Run a quarterly purple team exercise against the MDR specifically to measure response quality; vendor service reviews are not a substitute. Maintain an internal SOC lead who owns the MDR relationship and has the authority to terminate the contract; teams that delegate vendor management to procurement lose leverage and quality drifts.

The bottom line

A SOC is an organizational system, not a tooling problem. Tier structure, hiring criteria, onboarding design, metrics, and attrition drivers are interconnected; fixing one without the others produces marginal improvement at best. The leaders who build durable SOCs treat the team as a product with users, the analysts, and design accordingly.

Start with the most common failure modes: poorly tuned detections driving alert fatigue, hiring criteria that select for the wrong signals, onboarding that produces nominal but not real readiness, and a metrics culture that rewards velocity over judgment. Address those four and your SOC will outperform peers running the same tools at twice the headcount.

Frequently asked questions

How small can a SOC be and still be effective?

Below six analysts you cannot sustain 24x7 coverage with reasonable shift patterns and need either MDR augmentation for off-hours or a business-hours-plus-on-call model with documented response time expectations. Below three analysts, a SOC is essentially one shift of coverage plus the leader; this works for small organizations with limited attack surface but should be sized up before scaling the rest of the security program.

Should I require certifications for L1 hires?

No, with the caveat that a relevant certification is a positive signal among other signals. Requiring certifications as a hard gate narrows your pipeline and selects for credential-collectors over investigators. Hire on demonstrated investigation skill and writing ability; sponsor the certifications post-hire if the role requires them for compliance or career progression.

How do I justify L3 headcount to leadership?

Frame L3 as detection capability development rather than overhead. Track detection coverage contributed, rules tuned, alert volume reduced through tuning, and threat hunting findings that led to detection improvements. Present these alongside MTTD and false positive rate trends. L3 work compounds; one detection engineer can multiply the effectiveness of ten triage analysts, and that leverage is what justifies the role.

What is the right ratio of L1 to L2 to L3?

Roughly 4:2:1 in a mature SOC of 14 or more analysts, though it varies by alert volume and automation maturity. Highly automated SOCs trend toward fewer L1 and more L2 and L3 because triage is increasingly machine work; less automated SOCs require more L1 to absorb the volume. Track the ratio quarterly and adjust as automation reduces L1 load.

How do I handle an underperforming analyst without triggering attrition contagion?

Address performance early, specifically, and privately. Use the metrics conversation as the entry point: here is the pattern, here is what good looks like, here is the support plan, here is the timeline. Most underperformance is fixable with clearer expectations and targeted coaching; the cases that are not should be resolved decisively rather than allowed to fester. The most damaging attrition driver is not the departure of a weak performer but the visible tolerance of one.

Sources & references

  1. SANS SOC Survey
  2. MITRE 11 Strategies of a World-Class SOC
  3. Google SRE Workbook
  4. Phil Venables on Security Operations

Free resources

25
Free download

Critical CVE Reference Card 2025–2026

25 actively exploited vulnerabilities with CVSS scores, exploit status, and patch availability. Print it, pin it, share it with your SOC team.

No spam. Unsubscribe anytime.

Free download

Ransomware Incident Response Playbook

Step-by-step 24-hour IR checklist covering detection, containment, eradication, and recovery. Built for SOC teams, IR leads, and CISOs.

No spam. Unsubscribe anytime.

Free newsletter

Get threat intel before your inbox does.

50,000+ security professionals read Decryption Digest for early warnings on zero-days, ransomware, and nation-state campaigns. Free, daily, no spam.

Unsubscribe anytime. We never sell your data.

Eric Bang
Author

Founder & Cybersecurity Evangelist, Decryption Digest

Cybersecurity professional with expertise in threat intelligence, vulnerability research, and enterprise security. Covers zero-days, ransomware, and nation-state operations for 50,000+ security professionals every morning.

Black Hat Giveaway

Win a $2,495 Black Hat pass.

Full-access to Black Hat USA 2026 in Las Vegas. Subscribe free to enter.

Joins Decryption Digest daily briefing. Unsubscribe anytime.

Giveaway: Black Hat USA 2026 Full-Access Pass ($2,495 value)

Details →
Daily Briefing

Subscribe to enter the giveaway

Every subscriber is automatically entered. You also get daily threat intel every morning: zero-days, ransomware, and nation-state campaigns. Free. No spam.

Already subscribed? You're already entered.

Giveaway

Win a $2,495 Black Hat USA 2026 pass.