Win Probability Analysis for US Contracts: Guide

If your team bids on too many low-fit contracts, you waste time and money. I’d boil this guide down to one idea: score each pursuit the same way, use proof for every score, and tie that score to a clear bid or no-bid call.

Here’s the short version:

  • I use PWin to estimate how likely a deal is to close
  • I score core areas like:
    • technical fit
    • compliance
    • past performance
    • customer ties
    • price
    • delivery capacity
    • risk
  • I put each pursuit into a decision band, such as:
    • pursue
    • pursue with conditions
    • watch / no-bid
    • no-bid
  • I update the score when facts change, like a new RFP, amendment, Q&A, or competitor move
  • I check past scores against actual wins and losses to see if the model is off

A few numbers matter right away:

  • A weighted score of 4.0–5.0 usually means go
  • 3.0–3.9 means go only if gaps can be fixed
  • 2.0–2.9 often means stop unless there’s a strong business reason
  • 0–1.9 means no-bid
  • If compliance is 2 or lower, that’s usually a stop
  • Federal win rates are often low, so better pursuit selection can protect bid spend in a big way

What I like about this approach is that it replaces loose opinions with a written model. That makes it easier for sales, capture, proposal, pricing, and delivery teams to use the same logic instead of arguing from instinct.

Narwin.ai fits into that process by pulling in public bid data, scoring early-fit signals, reading RFPs, and keeping score history in one place. So the score becomes part of day-to-day pursuit work, not just a one-off spreadsheet.

That’s the whole point of the article: use a simple scoring model, back it with proof, review it as a team, and compare it to actual contract results over time.

US Contract Win Probability Scoring Model: Score Bands & Decision Guide

US Contract Win Probability Scoring Model: Score Bands & Decision Guide

Build a practical scoring model and weighting framework

A good scoring model doesn’t need to be complicated. It needs to be consistent.

Core factors, weights, and scoring bands

Start with seven factors that line up with how federal, state, and local buyers usually assess U.S. contract pursuits. Score each factor from 0 to 5, apply the weight, and calculate a weighted average on a 0–5 scale.

Factor Example Weight What It Measures
Technical Fit 25% How closely your solution matches the SOW, requirements, and evaluation criteria
Compliance Readiness 15% Ability to meet mandatory criteria, set-aside status, and security requirements
Past Performance 15% Relevant contract history, CPARS ratings, and references with comparable scope
Customer Relationship Strength 15% Incumbent status, stakeholder access, and understanding of agency priorities
Price Competitiveness 15% Your expected price versus budget benchmarks and competitor norms
Delivery Capacity 10% Staffing, subcontractor network, and operational ability to execute
Overall Risk 5% Legal, financial, and protest exposure that could threaten performance

To make scoring repeatable, use clear anchors for each number. A 5 means clearly superior, like multiple Exceptional CPARS ratings on directly comparable federal work. A 3 means acceptable but not distinctive. A 0 means a disqualifying gap, such as a missing certification or ineligible set-aside status.

The weights should shift based on the pursuit. For a lowest-price technically acceptable (LPTA) federal bid, put more weight on pricing and compliance. For a complex IT modernization, lean harder on technical fit and past performance. For a city construction contract, delivery capacity and risk often matter more than they would in a consulting pursuit.

Set go/no-go thresholds and action ranges

Next, turn the weighted score into a decision band so teams can act the same way every time.

Score Range Decision What It Means in Practice
4.0–5.0 Pursue Strong position; commit full capture and proposal resources
3.0–3.9 Pursue with conditions Proceed only if specific gaps can be closed; requires a mitigation plan and capture leader sign-off
2.0–2.9 Watch / no-bid Usually stop the pursuit unless there is a compelling strategic reason; require leadership review before moving ahead
0–1.9 No-bid Stop pursuit and redirect resources to higher-fit opportunities

Add two hard-stop rules. If compliance readiness scores 2 or below, treat it as no-bid unless the gap is fully closed before proposal submission. If pricing is materially above budget or award benchmarks, flag the pursuit as pursue with conditions and require a pricing review before proposal kickoff.

Once the model is in place, the next move is simple: define the internal and external data behind each factor.

Use the right inputs and data sources

Once you’ve set the core factors, each one needs an input you can actually trace back to a source.

A win probability score is only as good as the data behind it. If the inputs are shaky, the score will be too. And that leads to bad calls and wasted bid spend.

Common scoring inputs for US contract pursuits

Each factor in your scoring model should tie to something concrete and observable. Pull evaluation criteria straight from Sections L and M, then weight the model based on the priorities stated in the solicitation.

SOW fit looks at how much of the statement of work your team can handle with current capabilities, versus what would need subcontractors or added hiring. Mandatory compliance should be treated as binary: either you meet the certifications, clearances, set-aside status, and SAM.gov registration rules, or you don’t. If you can’t meet a required clearance or set-aside status, that’s a stop.

Scoring Input Strong Signal Moderate Signal Weak Signal
Evaluation criteria alignment All rated factors addressed with clear discriminators mapped to Section M Most requirements addressed; few differentiators Key evaluation areas missing or minimally addressed
SOW fit Near-identical scope at similar dollar value within last 3–5 years Related experience with partial task overlap Little relevant experience; unproven capability
Mandatory compliance All certifications, clearances, and registrations confirmed Minor gaps closable before due date Critical compliance gap with no fix before due date
Incumbent advantage You hold the current contract with strong CPARS and documented agency satisfaction Competitor incumbent under performance pressure Well-performing incumbent with long-term agency relationships
Past performance Multiple highly relevant references with excellent ratings at same agency or similar mission Limited relevance or mixed ratings; related domain No relevant references or negative performance history
Price competitiveness Target price informed by historical award data and market benchmarks Price based on internal estimates with limited market validation Price materially above or below historical norms without justification
Staffing and schedule Named key personnel largely on board; proven schedule performance Mix of named and contingent staff; plausible schedule High dependence on contingent hires; unrealistic transition timeline

Internal and external data sources that improve score quality

Scores depend on both internal and external data.

On the internal side, start with your CRM notes, capture plans, prior proposals, win/loss history, debrief feedback, and pricing records. Debrief notes are often overlooked. When you tag feedback to specific scoring dimensions – like "weak staffing plan" or "strong understanding of mission" – future teams can score similar pursuits with more confidence instead of rebuilding the logic every time.

Your pricing database should include:

  • Labor category rates
  • Indirect rates
  • Fee assumptions
  • Final submitted prices
  • Any award prices you can access

That gives you a way to compare future bids against what actually won, not just what your team expected.

On the external side, SAM.gov is the starting point for every federal pursuit. Sections L and M lay out the evaluation structure. Amendments show shifts in scope, pricing assumptions, and compliance rules. Q&A documents often hint at what the agency cares about most.

Federal award data from FPDS and USAspending shows winning contractors, award amounts, contract vehicles, and period of performance. That helps you judge price competitiveness based on actual award outcomes, not internal guesswork. Agency procurement history also shows whether a buyer tends to stick with incumbents or switch vendors. That has a direct effect on how you score incumbent advantage and customer relationship strength.

Narwin.ai aggregates many of these public sources – including SAM.gov and city-level portals – into a single feed, with AI-driven competitive intelligence and predicted win scores. Aggregation reduces manual tracking of amendments and award history.

Use these sources to fill in the model, then review and check scores against live pursuits.

Review, validate, and use scores in live pursuits

Once the model and inputs are set, the next step is deciding how teams review scores and act on them during live pursuits.

A score has no use if it just sits in a spreadsheet. It starts to matter when people review it, push back on it, revise it, and approve it through a clear process.

Review steps that keep scoring consistent

The review flow matters just as much as the model. Usually, the capture manager or opportunity owner scores first, using the agreed model and scoring bands. After that, capture, proposal, pricing, delivery, and contracts/legal check the assumptions.

Each factor should have a clear owner. For example:

  • BD owns relationship strength
  • Delivery/technical leads own solution fit
  • Pricing owns competitiveness

If a score changes, the team should document the evidence behind that change.

Before the review meeting, send the draft score and factor notes 24–48 hours in advance. That simple step gives reviewers time to think through the details instead of reacting in the room.

Outlier scores need the toughest review. For any score above 4.0, the team should test a few things: Has the customer shown concrete signs of preference? Is the target price assumption supported in this agency setting? Are competitor strengths being discounted too much?

For scores below 2.0, the conversation shifts. Has the team missed any teaming paths? Is the opportunity still important from a business standpoint despite the low score? Were the mandatory requirements and evaluation criteria read the right way?

Not every disagreement needs to end in forced alignment. If people see things differently, capture those views in the meeting notes.

Approval checkpoints should line up with major spend decisions, such as bid/no-bid, solution development, and proposal kickoff. Many mature federal contractors tie this process to formal gate reviews:

  • Gate 1 for qualification
  • Gate 2 for pursuit
  • Gate 3 for bid decision

At each gate, teams should log a time-stamped version of the score.

How teams apply scores across capture and proposal work

Scores should shape where time, budget, and people go.

A 4.0–5.0 score usually earns full capture investment. That can include deeper customer discovery, senior relationship support, and dedicated proposal staff.

A 2.0–3.9 score usually gets selective investment. The team stays engaged, but with more control over spend.

A score below 2.0 is usually a no-bid, unless there is a documented business reason to stay in.

The same logic applies to teaming. If an opportunity scores low on past performance or technical capability, that can trigger a search for partners with stronger credentials in the target agency. If the score is high, the prime may decide to lead instead of subcontract.

Pricing moves with confidence too. If the team has strong differentiation and clear customer preference, it may support a value-based price. If positioning is weak and competition is heavy, the pricing approach usually needs to be more aggressive.

Scores also change over time. They are not fixed. Teams should define trigger events that require re-scoring, such as:

  • a final RFP released after a draft
  • amendments that change scope or evaluation criteria
  • Q&A responses that clarify agency intent
  • new competitors entering the field
  • budget or procurement timeline changes

When one of those events happens, the capture manager should run a focused re-score on the affected factors and log the new version with a note that explains what changed and why.

Those same trigger points also build the dataset needed to test whether the model is working.

Validate the model against actual outcomes

Back-test recorded pursuits against actual wins and losses to check calibration.

Start with hit rates by score band. If high-scoring opportunities are winning far less often than expected, the model is too confident and needs recalibration. Then look at factor weights. Some factors may be getting too much influence, while others may not be getting enough, based on how they line up with actual outcomes.

If a factor shows weak or inconsistent correlation with wins and losses, reduce it or remove it. If the model is too blunt, add more detail. One good example is separating full-and-open competitions from task orders under established IDIQs. Those two cases have very different win patterns, so lumping them together can distort the score.

Run a formal model review at least once a year so scoring stays aligned with the way procurement practices and competitive conditions are changing.

A shared workspace makes this much easier to manage.

Narwin.ai supports that loop by keeping win scores, score rationale, and RFP analysis in one workspace for comparison over time.

Operationalize win probability analysis with Narwin.ai and key takeaways

Narwin.ai

Where Narwin.ai fits in the workflow

Once the model is defined and validated, the next move is to put it to work in day-to-day pursuit activity.

Narwin.ai brings opportunity discovery, scoring, and proposal drafting into one workflow. At the opportunity discovery stage, it monitors public procurement sources such as SAM.gov and local and state portals. It surfaces relevant opportunities and assigns a preliminary match score before your team spends time on a deep manual review. That score is based on your profile, certifications, capabilities, past performance, and eligibility.

After opportunity screening, the platform helps teams move from qualification to response drafting. Narwin.ai’s RFP analysis engine ingests solicitation documents and pulls out requirements, evaluation criteria, compliance clauses, and key dates automatically. Each requirement is then scored for compliance, quality, evidence, and clarity, so teams can spot strengths and gaps fast.

There’s also a Risks tab that flags blockers like tight timelines, bonding requirements, and vague scope language, along with suggested mitigations. From there, the platform generates a color-coded win score with the reasoning attached. If the team decides to pursue, the AI Writer uses extracted requirements and your indexed company knowledge to produce tailored first-draft proposal sections.

The whole workflow runs in one platform and connects with Google Drive, Slack, CRM, and ERP systems. Full setup details and workflow guides are available at narwin.ai/docs.

The main point isn’t just the score. It’s the fact that the same model can follow the opportunity from discovery through the award decision. That turns scoring into part of the pursuit workflow instead of a separate reporting task.

Conclusion: The essentials of a reliable win probability process

A reliable win probability process uses:

  • A fixed scoring model
  • Traceable evidence
  • Cross-functional review
  • Score updates when facts change
  • Validation against actual outcomes over time

Narwin.ai supports that process by centralizing scoring, evidence, alerts, and score updates in one workflow.

FAQs

How do I set weights for different contract types?

Pick five to ten factors that have the biggest impact on your odds of winning the deal. Common ones include customer relationships, past performance, technical strength, pricing, and compliance.

Then assign each factor a percentage weight based on how much it matters for that specific opportunity. The weights should add up to 100%.

For example:

  • In a recompete, customer relationships may matter more, so that factor should carry more weight.
  • In new business, technical strength and pricing often matter more, so those factors should get a larger share.

Once your factors and weights are set, calculate pWin with a weighted formula. A simple approach is to score each factor, multiply that score by its weight, and then add the results together.

pWin = Σ (factor score × factor weight)

This gives you a more grounded pWin estimate based on the parts of the deal that matter most.

What evidence should support each PWin score?

Each PWin or win probability score needs to rest on plain, checkable proof, not gut feel.

That proof should come from direct citations in the RFP, especially Section L and Section M, along with source material like CPARS reports, past performance references, and project history. If a score says your team is strong in one area, the backup should show exactly why.

Other support can come from customer call notes, meeting summaries, and internal gap analyses for items like certifications or clearance levels. For manual scoring, use standardized rubrics that require a citation for every rating so the score stays tied to facts instead of opinion.

How often should we re-score an opportunity?

Re-score an opportunity whenever new information comes in, like updated proposal data, revised requirements, competitor intel, solicitation amendments, or Q&A responses. Even small shifts can change your win odds and the way you should play the pursuit.

It also helps to review win predictions every quarter against actual results. That gives you a simple reality check and helps you fine-tune your weighting system over time. Tools like Narwin.ai can make it easier to keep scores current without a lot of manual work.

Related Blog Posts