I have seen growing interest among funders in translating their investments into an impact value that can strengthen due diligence or inform strategy. For some that means a cost-effectiveness number or a benefit-cost ratio. For others it is a social return on investment, or an estimate of the total downstream impact a portfolio generates.

For funders considering this, the reality is that there is no standard model to recommend. Every approach carries tradeoffs, and the right one depends on how the funder works.

Rigor is absolutely necessary, but funders often overestimate how much methodological rigor alone determines whether the system succeeds. A framework that takes a team two months to value a single investment can be methodologically excellent and still fail at an organization that needs to decide within three weeks and funds 100 grants a year.

In practice, the hardest questions are often operational and institutional.

Who is allowed to set key assumptions? How are disagreements handled? What evidence standards are realistic given the pipeline? Can the approach be maintained year after year without specialist support?

A valuation system that does not fit how the funder works will be quietly abandoned, no matter how defensible its numbers.

Through my work with foundations, impact investors, and philanthropy-serving organizations, I've found that the success or failure of a valuation system is usually determined by five design clusters:

  • A. Operational fit
  • B. Analytical defensibility
  • C. Portfolio strategy and aggregation
  • D. External credibility
  • E. Governance and institutional design

Two of them, operational fit and governance, are the ones funders most often overlook, because attention flows to the methodology long before anyone asks how it will be run.

Companion tool
Model Design Tradeoff Map
An interactive map of where the leading impact valuation methodologies sit across the five design clusters below.
Companion resource
Impact modeling frameworks for funders
A reference directory of impact valuation methods and their primary public sources.

Cluster A: Operational fit (can you implement it?)

The most overlooked cluster, and often the binding one. A method can be analytically perfect and still be unusable at the volume and staffing a real portfolio has. A method too slow or too demanding to keep up with does not get improved, it gets abandoned, as staff fall back on the tools they already trust.

A1. Implementation feasibility

Can existing staff (program officers or impact officers) run the models themselves or does the method require dedicated economists, external consultants, or proprietary tools? Spreadsheet-native methods such as income-based ROI estimators sit at one end. Multi-multiplier monetization methods that need specialist staff or vendors sit at the other. This is a consistency question as much as a cost one, because outsourced models are harder to apply uniformly across a portfolio.

A2. Speed and volume

Time per deal matters, but volume at scale matters more. A method that takes four weeks to analyze per deal works for 12 deals a year and breaks at 100. Track time to first-pass impact estimate, time to refined estimate, and time to portfolio rollup separately, because a method can be fast at one and slow at another. A program officer may have to get back to a potential grantee within a couple of weeks — will the impact model slow that down?

A3. Applicant data burden

What must the applicant provide? Rough impact projections only, or detailed theory-of-change maps, baseline and control data, and full impact modeling? Cost detail belongs here too. A method can take a grantee's headline figure at face value, a million dollars to reach a thousand people, or require itemized cost accounting behind it. More detail makes the denominator harder to inflate and raises the burden on the applicant, the same tradeoff this cluster keeps surfacing in a different place. High burden filters out early-stage and low-capacity grantees regardless of their impact, and biases the portfolio toward larger, better-resourced organizations.

A4. Reusability and updatability

Does the method produce numbers you can update as evidence arrives, or does each assessment stand alone in isolation? Methods built for annual refresh stay useful as living tools. Methods that produce one-time snapshots quickly stop matching the reality on the ground.

A5. Tooling and template availability

Is there an existing published template, spreadsheet, or tool, or do you build the infrastructure before you can start? A ready easy-to-use tool lowers adoption cost. The absence of one is a hidden project that lands on your team before the first grant is evaluated.

Cluster B: Analytical defensibility (will the numbers hold up?)

This is the cluster the published literature argues about most, and the one funders assume matters most. It does matter, but defensible turns out to be two questions stacked on top of each other, and the field usually only debates the second. The first is how value is expressed: the unit it is in, the structure it takes, and what it is measured against. The second question is whether that number survives once its shape is fixed, and that is where the rest of the cluster lives.

How value is expressed

B1. Output unit

Some frameworks keep impact in a non-monetary unit, a Quality Adjusted Life Year (QALY), a WELLBY, a standard deviation of test score, while others convert everything to dollars. A non-monetary unit keeps the underlying evidence visible and sidesteps a contested conversion factor. Monetization buys a common denominator and speaks to a finance audience, at the cost of pricing outcomes that resist a price, such as the cost per QALY. Another option is keeping the output unit only to real income gained or costs averted which reduces the need to focus on monetizing welfare outcomes.

B2. Output structure

The field has settled language for this one.

Cost-effectiveness analysis (CEA) reports cost per unit of a single outcome, such as cost per life saved and cost per job created. It can be harder to do cross-sector comparisons if the outcome is specific to one impact sector.

Cost-benefit analysis (CBA) converts every outcome to dollars and sums them, which buys cross-sector comparability at the price of pricing things that resist a price, although most CBA models only focus on monetary benefits such as income and wealth outcomes. The output is then in a ratio of benefits divided by costs which make comparability easier to see for decision makers.

Social return on investment (SROI) is structurally a cost-benefit analysis whose benefits are deliberately non-market. SROI translates non-monetary outcomes into financial proxies such as the value of social connectedness, environmental improvements, or improved mental health.

Comparability needs will often drive how value is expressed. A funder working inside a single sector can live with cost-effectiveness and never monetize. A funder running a cross-sector portfolio cannot compare cost per outcome across strategies, which forces cost-benefit or social return on investment, which then forces the monetization question.

B3. Reference point and headline metric

A framework can report an absolute figure, 10 dollars of value per dollar spent, or a multiple over a fixed comparator, six times a cash benchmark. The frameworks that monetize do not converge on one structure. GiveWell expresses cost-effectiveness as a multiple over a cash benchmark, the consumption gain from giving cash to people in poverty, with a funding bar at six times that benchmark as of May 2026, and it states plainly that the model exists to test whether a grant clears the bar rather than to rank finely above it. Coefficient Giving also reports a ratio, an SROI of roughly 2,000x, but its denominator is a constructed welfare unit rather than a comparator program.

That structure changes what the funder has to defend. A benchmark multiple asks you to stand behind a relative claim, that a grant beats a known reference by a known factor, rather than an absolute value-of-life figure. That is why a benchmark approach can tolerate cruder absolute numbers than a pure cost-benefit model.

Whether it survives scrutiny

B4. Evidence standard required

What is the minimum standard of evidence needed to produce the impact valuation estimates? Does the framework demand RCT or quasi-experimental evidence, tolerate tiered and mixed evidence, or accept a theory of change with stated assumptions? A hard RCT gate gives the strongest causal claim and filters out exactly the early-stage and low-capacity grantees who often hold the highest impact potential per dollar. Evidence flexibility widens the pipeline and shifts the defensibility burden onto stated discounts and analyst judgment. Some funders treat this as a preference. Others treat it as a hard line in the sand, and for them it is one of the most important design choices they make.

B5. Counterfactual and attribution

How are deadweight (what would have happened anyway), attribution (share of outcome caused by others), displacement (benefits that displace value elsewhere), and drop-off (decline in outcomes over time) handled? Does the model only rely on causal research to estimate the counterfactual, or will it use solely an analyst's judgment to estimate deadweight? Empirical, evidence-driven counterfactuals are more reliable but more demanding to run. Judgment-based discounts are quicker but raise more rigor-related questions. The same logic runs on the denominator. When a grant is co-funded, attribution governs cost as well as benefit, because claiming the full result against your share of the money inflates the ratio by exactly the fraction you did not pay for. A complete denominator counts total resources mobilized, co-funder money, grantee contributions, and leveraged public funds, not only your own check.

B6. Marginal additionality

Additionality of the next dollar fails in two ways. The first is crowding-out, when another funder would have backed the grant anyway, so your money frees up theirs rather than creating new impact, and its mirror image is crowd-in, when your funding pulls in other capital and does more than your dollars alone. The second is saturation, when a grantee is near the limit of what it can productively spend, so the next dollar buys far less than the average dollar already has.

A method that ignores both misreads exactly the grants where allocation is hardest. It overvalues a saturated grantee whose average impact still looks strong but whose room for more funding is gone, and it undervalues a catalytic or matching grant whose whole purpose is to move other people's money. Average impact tells you whether a grant was worth funding. Marginal impact tells you whether it is worth funding more.

Cluster C: Portfolio strategy and aggregation (are the numbers comparable?)

Almost every framework was built to value a single grant. Few provide guidance on how grants can be summed into a defensible portfolio total. The moment you want one headline number across a diverse portfolio, the challenge becomes aggregation.

C1. Outcome breadth and scope

Does the method carry one outcome or many, and how far down the causal chain does it reach? Single-dimension methods (income, emissions) are simplest and exclude what they cannot measure. Composite monetized methods carry several dimensions at the cost of commitments about cross-dimension trade-offs. Scope is the second half of this choice: most methods handle direct beneficiary impact well, household and community spillovers imperfectly, and ecosystem or systemic effects barely at all, so a funder backing R&D or policy needs a method built for the wider scope.

C2. Cross-country comparability

When a portfolio spans countries, a framework that sums absolute dollar gains rates grants in higher-income places higher, because the same intervention moves more money where wages are higher. A 30 percent income gain on a US wage is a far larger dollar figure than the same 30 percent gain on a Kenyan one, even when the two grants do equal good for the people they reach. Cost-of-living differences are why those larger figures mislead, since a dollar of gain buys less real welfare in a high-cost country than in a low-cost one. A framework that ignores this is not neutral. It carries a built-in tilt toward rich-country grants that no funder chose on purpose.

Two responses exist, and the choice sits at the portfolio level so every grant is scored on the same basis. One is to convert gains to a purchasing-power basis, so nominal dollars reflect comparable real consumption across countries. The other is to measure income gains relative to baseline rather than in absolute dollars, which removes the income-level effect at the source.

The GitLab Foundation takes the second route and publishes both numbers, the North Star ROI measures absolute lifetime income gains and the Relative Income Change model (DIL, Double Income for Life) measures the gain as a percentage of baseline, so raising a low income by a given share counts as much as the same share on a high one. They built the relative version because they compare grantees across regions where average incomes vary widely, and an applicant can clear either threshold to qualify for funding.

C3. Aggregation rule setting

Whether you can roll up is the easy question. Whether the rollup is honest is the design choice. Summing grant-level impact double-counts anyone reached by more than one grant, so a true count of individuals reached has to de-duplicate the students touched by both a direct-to-student program and a teacher-training program. Adding effects on the same people is its own trap, because stacked interventions rarely sum. For example, students who receive both edtech and a better-trained teacher will not show the learning gain of the first plus the learning gain of the second, since the two interact rather than add.

Portfolio rollups implicitly assume an aggregation rule. The rule may be additive, diminishing, or synergistic. Most systems never state which rule they are using.

C4. Out-of-bounds

A mature system decides where its numbers stop, and it does so in two places. The first is what not to model at all, because a breakthrough bet, an advocacy play, or early field-building often may have no defensible number yet, and forcing one trades an honest blank for false precision, so treat those qualitatively and say why. The second is what not to compare, because two grants can both carry credible numbers and still not share an axis when their denominators are different kinds of things, like cost per life saved against cost per job created. A system that scores everything eventually gets quoted on grants it was never built to judge.

Cluster D: External credibility (how does it land outside?)

The outward face of the system. Whether the method is good and whether outsiders accept it are different questions, and a method can be strong on one and weak on the other. This cluster is about peers, boards, auditors, and donors. Whether your own staff will run it is a separate test, and it lives in the next Cluster E, Governance.

D1. Field acceptance

Is the method recognized by peers, co-funders, and board members? A widely used standard requires less defense. A method no one else uses requires more defense regardless of its quality, which is a real cost even when the method is sound. For example, SROI has a published framework, the IFC and GIIN manage the Operating Principles for Impact Management, and the Impact Management Project's Five Dimensions of Impact is known to be a standard framework.

D2. Transparency and reproducibility

Are assumptions, parameters, and ideally the model itself published? The most transparent version publishes the full model down to the spreadsheet, e.g. GiveWell's public cost-effectiveness models and GitLab Foundation's public models, so any reviewer can trace a headline figure to its inputs. For funders who expect methodology challenges from sophisticated donors, published-and-reproducible is worth the exposure.

D3. Independent assurance

Can outputs be externally verified? Some methods have formal assurance pathways and verification standards with public disclosure of the result. For institutional funders accountable to boards or LPs, an assurance pathway is sometimes a hard requirement rather than a preference.

Cluster E: Governance and institutional design (will it survive inside your organization?)

This is where many valuation systems fail. The clusters above describe the properties of a method. Governance determines whether the method gets trusted and used once real funding decisions are attached to it.

The core tension is straightforward. Leadership often wants a common unit that enables comparability and a defensible answer for the board. Program officers, especially sector specialists, may experience the same framework as a flattening of work they know to be complex. A framework adopted by leadership without program-officer buy-in becomes a compliance exercise. One adopted by program officers without leadership agreement on how it should be used produces a number nobody acts on. Impact valuation has a lot of misconceptions and clear communication is critical to bring staff along.

This framework does not assume that funders alone define value. Some institutions may choose to share or delegate decisions about outcomes, assumptions, weights, or funding priorities to grantees, community members, participatory grantmaking bodies, or other stakeholders. Those choices are themselves governance decisions and can be evaluated using the same design clusters described here.

Governance also covers how disagreement gets handled once the assumptions are fixed, which the items below take up. A perfect method without the right governance and institutional fit gets decommissioned within a few quarters.

E1. Parameter ownership

A model runs on two kinds of inputs, and parameter ownership is the question of who is allowed to set each. Systemic constants such as the discount rate, effect persistence, earnings conversion, and default attribution apply across every grant in the portfolio. Grant-specific inputs such as the number of people reached, baseline income, cost, and the assumed effect size apply to one grant only. The ownership question runs along both, not just the constants, because each carries a different risk when the wrong person holds it.

For the constants, the choice is whether the program officer building the grant's case sets them, or a central monitoring, learning, and evaluation (MLE) team locks them and refreshes them on a fixed cadence. Grant-specific inputs can be verifiable facts, like cost and the number of people reached, which the program officer or grantee supplies and the MLE team can check against records. Others are grant-level judgments, like the effect size assumed for an untested program or the attribution assigned in this particular context, and those carry the same drift risk as the constants. If the program officer who wants the grant funded sets the grant-level judgments, there is risk of gaming the model.

The constants could be locked centrally. The factual inputs are supplied by the program officer or grantee and validated against records by the MLE team. The grant-level judgments are the contested middle, and one approach is that the program officer proposes them from contextual knowledge while the MLE team sets the final value and applies a consistent rule across grants, which keeps the program officer's knowledge in the model without letting the person with the incentive set their own number.

Where a funder lands on those judgments is the live tension. Push them toward program officers and the model tracks each grant more closely but drifts toward convenient values. Push them toward MLE and the portfolio stays consistent, but program officers can come to see the model as something applied to their grants rather than built with them, which is where adoption starts to fade.

E2. Stakeholder voice and decision rights

E1 asked which internal actor sets each input. This asks whether that authority stays inside the institution at all, and how far it travels when it does not. At one end the funder or its analysts set the values and the people closest to the outcome are not consulted. Consultative approaches gather beneficiary and grantee input to sharpen the funder's judgment without moving the decision. Participatory approaches cede the decision itself, letting community members or a participatory grantmaking body set priorities, weights, or the funding call.

Two channels matter here, and equity-focused funders weigh both. The first is beneficiary voice over what counts as valuable, rather than an analyst assigning the weights for them. The second is grantee feedback, whether the organizations being funded can question the model and whether it changes in response, since a model that improves on grantee input stays trusted and one that ignores it gets worked around. The further this authority moves toward delegation, the more it is in tension with operational fit, A2, because wider participation is slower to run.

E3. Dissent discipline

A model that lets any senior person quietly move the base assumptions has no valid assumptions. The alternative is a logged tracking of where analysts disagree with the key assumptions. The objector's edits sit alongside the base case rather than replacing it, the base stays stable, and the distribution of dissents becomes its own portfolio records worth reading. It creates a number that means the same thing each time it is quoted.

E4. Use boundaries on how the model is used

Some uses of the impact models must be fenced off. Do not use an ex-ante estimate to evaluate staff performance after the fact, because it can corrupt the model input assumptions. Do not let small ROI differences drive prioritization, because they sit inside the uncertainty bands. Do not use a low number to kill a strategically important grant whose value the framing was never built to capture. A number allowed to do anything gets weaponized, and fencing protects the trust that keeps it in use.

E5. Internal adoption

External recognition is whether peers and boards respect the method. Internal adoption is whether your own program officers will run it, and it is where adoption usually dies. Resistance is rarely about the math, it is about scope, ownership, and whether the method fits alongside the models the program teams already trust. The executive version of this test is whether the method gives senior leadership one headline figure a CFO will stand behind, with a clean audit trail back to grantee-verified outcomes.

E6. Composability with program-team models

Program teams often already have trusted models for their own outcomes. A portfolio framework that tries to replace those gets rejected. One that sits on top, takes the program team's number as the numerator, and adds only the cost denominator and the cross-portfolio currency, gets adopted. The rule that follows is uncomfortable for a central measurement function and correct anyway: program-team estimates take precedence over centrally generated ones for the outcomes the program team owns.

The tradeoffs are real

The five clusters are partially anti-correlated, so no single framework wins on all of them. Methods that score highest on external credibility (D), with formal assurance and consultation requirements, tend to be the slowest to run and so score lower on operational fit (A). Methods that score highest on operational fit, the self-serve spreadsheets a program officer can run in days, tend to accept applicant-supplied projections and so score lower on analytical defensibility (B). Methods that score highest on portfolio strategy (C) require a common unit that program officers may experience as overly reductive, which raises the governance cost in (E).

The result is not a ranking. It is a tradeoff space, and the right framework is the one whose strengths line up with what your foundation needs and whose weaknesses fall where your decisions do not depend on them.

This is not an argument against standardization, only against standardizing the wrong layer. The case for a common approach is real. A nonprofit reporting to a dozen funders in a dozen formats carries a cost the field rarely counts, and shared structure would lower it. The mistake is to read that as a case for one shared valuation model, the part that turns outcomes into a number, because that is the layer where funders genuinely differ and the tradeoffs above genuinely bind.

What can and should converge is disclosure. Two funders can run different models and still report the same assumptions in the same way, the discount rate, the years of impact assumed, the evidence tier behind each effect size, the counterfactual and attribution adjustments, and the comparator the headline figure is measured against (IDInsight's Jeffrey McManus makes a similar argument). Standardizing those makes results legible across funders and lets a reviewer see why two numbers differ, without forcing every funder onto a model that fits none of them well. Converge on what you disclose. Match the model to what you decide with.

Companion tool
Model Design Tradeoff Map
An interactive map of where the leading impact valuation methodologies sit across the five design clusters described above.
Companion resource
Impact modeling frameworks for funders
A reference directory of impact valuation methods and their primary public sources.

The model comes after the strategy

An impact valuation model compares grants on one quantity, how much value each produces per dollar, and it is blind to everything that quantity leaves out. If you care about a particular population, geography, or income level, the model will not honor that unless it is explicitly built in. A cost-benefit analysis comparing a job training program for youth against an equivalent program for Indigenous communities will favor whichever shows the higher monetized outcome per dollar. The model does not see the population. It sees the number.

This is why valuation belongs after strategy, not before it. A funder focused on early-stage cancer research should define that strategy first and screen the pipeline down to the grants that fit it, then run the valuation model on the grants that remain. The impact model ranks options within a mandate. It does not choose the mandate.

So treat the model as one input in a wider due diligence process, not the whole of it. Score each grant on strategic fit and on fit with your focus populations as their own criteria, alongside the valuation number rather than inside it. Some models can carry distributional weight, and the relative-income approach in Cluster C2 is one example. But most published models leave equity out by default, so do not assume the number reflects priorities you never asked it to weigh. A grant may score lower on value per dollar and still be a strong match with your existing strategy focus.

Strategy gets you to a subset of grants that fit your mission. The model tells you which of those grants does the most good per dollar, which is the comparison that strategy alone can't make. It also highlights the exceptions. When you fund the lower-scoring grant because it serves a population you exist to serve, the model tells you what that choice costs in foregone measured impact, which makes the tradeoff visible rather than hidden.

The choice that actually matters

Choosing an impact valuation approach is a system design problem, not a framework choice. Getting the methodology right is necessary, and it is not sufficient. The harder part is whether the method fits what your team can run, what your board will accept, and how the resulting number gets owned once real money is attached. A funder who gets the methodology right and the feasibility and governance wrong ends up with a defensible number nobody uses.

The frameworks are tools. The judgment is in matching them to the mission.