## _wp05210

## Source details

**Canonical URL:** [_wp05210](https://www.imf.org/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2005/_wp05210.pdf)

## Other formats

- [Markdown version](/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2005/_wp05210.pdf.md)
- [Structured JSON version](/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2005/_wp05210.pdf.json)

---

### I. Introduction: scope and motivation
- Performance budgeting renewed interest since the 1990s across advanced, developing, and transitional nations.
- Reform drivers: perceived disappointing public sector performance and need for fiscal sustainability after late-1980s fiscal consolidations.
- Noted skeptical assessments:
  - Schick (2003, p. 83): “efforts to budget on the basis of performance almost always fail.”
  - U.S. Congressional Budget Office (CBO, 1993, p. x): “little evidence of the much-touted advances in performance-based budgeting in local, state, and foreign governments. Performance measures have a limited ability to influence the allocation of resources.”
- Primary aims:
  - Review empirical literature on government-wide performance budgeting systems since 1990, distinguishing quantitative analyses from case studies.
  - Examine empirical literature on sectoral performance budgeting systems, with emphasis on casemix hospital funding.
- Methodological caveats:
  - Performance budgeting is embedded in broader reforms, making attribution problematic.
  - Success depends on contextual factors (political system, political culture, fiscal environment).
  - The concept of performance budgeting is variably defined.

### II. Definition, scope, and objectives of performance budgeting
- Working definition:
  - Procedures/mechanisms intended to strengthen links between funds provided to public sector entities and their outcomes and/or outputs through use of formal performance information in resource allocation decision-making.
  - “Formal” performance information includes performance measures, cost measures, and effectiveness/efficiency assessments obtained through analytic tools.
- Core objectives:
  - Enhanced allocative efficiency.
  - Enhanced productive efficiency (technical efficiency and absence of waste/X-inefficiency).
- Scope:
  - Encompasses government-wide budget formulation, internal agency resource-allocation decisions, and sectoral/service-level output design choices.
  - Examples include program classification, outcome/output targets linked to funding, formal funding “contracts,” formula-based output funding, output-purchase budgeting, and casemix hospital funding.
- Three levels of allocative decision-making:
  - Government-wide budget decisions.
  - Resource allocation within agencies.
  - Service (output) design decisions affecting program effectiveness.

### Key procedures and mechanisms
- Program classification to facilitate prioritization and accountability.
- Setting outcome or output targets linked to funding (e.g., U.K. PSA; OECD survey: of 27 OECD nations responding, 12 asserted they set performance targets as part of the budget process, and in 9 of these the results were reported ex-post to the finance ministry).
- Formal agreements/contracts articulating funding-linked targets.
- Formula-based budget estimation: planned output volume × per-unit funding amount (examples: school funding by student course-years; U.S. Working Capital Funds).
- Formularized output-based supplementary funding for extra demand (example: Australia “workload agreements” where immigration department received additional funding per detainee/day).
- Output-purchase budgeting and casemix funding: payment for outputs delivered; DRG-based classification with per-output prices.

### Misconceptions and limits
- Reporting performance measures in budget documents is not sufficient; emphasis is on budgetary use of performance information.
- Not all systems reward/punish ex post performance; some link funding to future targets or non-budgetary sanctions.
- Performance budgeting is not synonymous with broader managing-for-results reforms; its distinguishing feature is budgetary use of performance information.
- Tying managers’ hands on input mix guarantees technical inefficiency; relaxing input controls is a core institutional ingredient.

### Why pursue performance budgeting?
- Motivations:
  - Counteract incrementalism and improve responsiveness of allocations to social needs.
  - Increase pressure on agencies to lift performance via budget-linked targets, prices, and agreements.
  - Achieve a “productivity dividend”: if productive efficiency rises, aggregate public expenditure should, ceteris paribus, be smaller, allowing lower tax burdens or funding new priorities.

### Empirical evidence: government-wide systems — overview and methodological issues
- Major methodological obstacle: absence of a practicable way to measure degree of allocative (in)efficiency for a given budget allocation; literature focuses on changes in allocation and subjective assessments rather than causal attribution.
- Empirical evidence organized under three headings: budgetary allocation, aggregate expenditure, and productive efficiency/program effectiveness.
- Findings emphasize limited, largely U.S.-focused, survey and case-study literature; quantitative expenditure-based analyses are scarce.

### Survey and case-study findings (selected exact figures and results)
- Melkers, Willoughby et al. (1999–2000 GASB survey):
  - 7.5 percent of state budget office officials believed performance measures were “very effective” or “effective” in changing budget appropriations.
  - 26.3 percent of state executive budget offices considered output and outcome measures to be either “important” or “very important” in budget appropriation decisions.
  - For state budget offices responses (36 states): percentages asserting performance measures were “very effective” or “effective” in changing budget appropriations: 7.5 percent (state budget offices), 24.2 percent (state agencies), 16.9 percent (city/county).
  - 26.3 percent of respondent state budget offices affirmed that output or outcome measures were either “very important” or “important” in decisions about budget appropriations to agencies.
  - 35.1 percent believed output or outcome measures were important or very important in agency “budget development.”
  - 28.1 percent believed these measures were important or very important in agency “budget execution.”
- Jordan and Hackbart (1997):
  - 29 out of 45 respondent states agreed that “achievement of performance standards affects budget recommendations in the Governor’s Executive Budget.”
  - 33 out of 46 claimed performance affects funding in the next fiscal year; 11 of the 46 claimed performance affects funding in the current fiscal year.
- Poister and Streib (1999) — city managers survey:
  - 60 percent of respondents from jurisdictions with “centralized, citywide performance measurement systems that incorporate most departments and programs” believed performance measures brought about moderate or substantial changes in city budget allocations.
  - 46.4 percent of those from cities with citywide performance measurement systems believed performance measures had a moderate or substantial impact in reducing the costs of services.
  - 71.6 percent believed performance measures had a moderate or substantial impact on the quality of services.
- GAO (2000) federal managers’ survey (2,510 usable responses):
  - 43 percent of respondents reported using performance measures to a “great” or “very great” extent in making resource allocation decisions.
  - Similar percentage (45 percent) reported using performance measures for priority setting.
- GASB case studies (17 jurisdictions):
  - In only 6 out of 17 jurisdictions did it appear performance measures were significantly used in and impacted resource allocation decisions.
  - Of eight jurisdictions not newcomers to performance budgeting (introduced no later than mid-1990s), six claimed substantial effect on budget allocations (Oregon, Texas, Iowa, Winston-Salem, Prince William County and Sunnyvale).

### Quantitative analyses and incrementalism literature (selected exact results)
- Reddick (2003a): rational budgeting (program budgeting, ZBB, biennial budgeting) associated with restraint of spending levels in only a minority of functional spending sectors.
- Reddick (2003a) on U.S. state budgets 1960–96: incrementalism had “relatively low explanatory power” for broad functional categories.
- Dezhbakhsh, Tohamy and Aranson (2003): U.S. federal nondefense expenditure 1946–1996, firmly rejecting incrementalism and identifying triggers for major shifts.
- Post-1990 literature does not support a strong form of incrementalism; incrementalism may coexist with punctuated reallocations.

### Aggregate expenditure impacts
- Reddick (2003a): states with program budgeting, ZBB or biennial budgeting more likely to have restrained aggregate expenditure after controls; causality not established.
- Brumby, Edmonds and Honeyfield (1996) on New Zealand: post-reform bold reprioritization decisions and reduced central government expenditure as a proportion of GDP; insufficient evidence to attribute causality solely to financial management reform.

### Productive efficiency and program effectiveness: evidence gap
- Very limited literature measuring improvements in output costs and outcomes attributable specifically to performance budgeting.
- Brumby, Edmonds, and Honeyfield (1996): unit cost data for four “process” outputs from three ministries showed data for three of four outputs suggesting significant productivity gains; attribution to performance budgeting versus broader reforms is unclear.
- Surveys suggest perceived improvements: 46.4 percent (cities) reported cost reduction effects; 71.6 percent (cities) reported quality improvements; GAO (1996–97) found 42 percent of federal managers believed GPRA improved programs to a “moderate or greater extent.” Yet these are self-reported perceptions.

### Sectoral performance budgeting — casemix (DRG) hospital funding: mechanisms and rationale
- DRG classifications group acute in-patient treatments into categories “relatively homogenous in respect of the resources used.”
- Prospective Payment Systems (PPS) and DRG pricing: hospitals receive fixed prices per DRG; difference between actual cost and DRG price represents loss or profit to hospital.
- International adoptions: Portugal (1990), Australia (from 1993), Norway (1997), Singapore (1997), United Kingdom (2004).
- Policy intent: create financial incentives for efficiency, increase treatment volumes, reduce waiting times, and reduce input controls over managers.

### Casemix empirical findings: cost containment and quality (selected exact results and studies)
- Coulam and Gaumer (1991) synthesis (to 1991): PPS had been highly successful in improving hospital efficiency and reducing Medicare cost growth; “none of the worst fears raised at the outset have been borne out by experience”; no evidence of significant deterioration of health outcomes to that time.
- Boccuti and Moon (2003): compared Medicare cost increases with private insurers and concluded: “Medicare has proved to be more successful than private insurance in controlling the growth rate of health care spending per enrollee.”
- Meltzer, Chung, and Basu (2002): one post-1991 U.S. paper finds some evidence of “skimping” in California.
- Portugal: Dismuke and Guimaraes (2002) found “no evidence” that case-mix payment had a pernicious effect on hospital mortality for a studied DRG.
- Portugal: Dismuke and Sena (1999) found casemix funding had a positive effect on productivity of diagnostic technologies in three common technologies.
- Victoria, Australia studies:
  - Young and Harris (1999): results “consistent with the expectation that casemix funding is providing hospitals with appropriate incentives for cost efficiency,” with methodological caveats.
  - Victorian Auditor-General (1998): official cost data showing dramatic cost reductions and changes in treatment patterns; Auditor-General also considered there had been an “overall decline in the quality of care” based on opinion surveys (methodologically questionable).
  - Brown and Lumley (1998): patient satisfaction surveys for maternity care showed a doubling of critical comments 3 months after casemix introduction; authors concluded “standards of care are being compromised in the current economic and policy environment.”
- General contemporary view summarized: “overall ...the system has worked pretty well in establishing better incentives without having a systemic adverse effect on quality.”

### Behavioral responses, risks, and mitigation under casemix
- Potential dysfunctional behaviors: “skimping,” “dumping,” “creaming” (cream-skimming).
- Mitigations:
  - Quality-oriented regulatory structures and monitoring regimes.
  - Payments for anomalous “outlier” patients.
  - Progressive refinement and disaggregation of DRG classifications.
- Role of professional altruism: socialized, value-based drivers and professional standards acted as buffers against purely financial incentives degrading quality; empirical quantification absent but considered important.

### Cross-sectoral generalizations and caveats
- Casemix success does not guarantee transferability; what works for hospitals may not work in education, police, or other health subsectors.
- Empirical evidence of perverse effects exists in other sectors (school education; labor market programs like JTPA).
- Altruistic motivation varies by sector; strong in hospitals where stakes of care failure are severe.
- Managing-for-results systems can complement altruistic motivations by communicating broader cost-effectiveness concerns to practitioners.

### Behavioral foundations: motivation, incentives, and public service motivation
- Organizational economics cautions high-powered extrinsic incentives under poor measurability, uncertainty, and multi-tasking can produce perverse responses; recommends greater reliance on lower-powered, “fuzzier” extrinsic incentives or fixed remuneration plus intensive monitoring (Gibbons, 1998, p. 120).
- Empirical evidence:
  - Public sector workers exhibit stronger “public service motivation” (Crewson, 1997, p. 515).
  - Excessive extrinsic incentives can “crowd out” intrinsic motivation (Osterloh and Frey, 2000).
  - The literature suggests optimal systems mix extrinsic incentives with support for intrinsic motivations; evidence exists that well-designed pay/performance can complement intrinsic motivation (Deckop, Mangel and Cirka, 1999; Taylor and Pierce, 1999).
- Recent theoretical and experimental contributions: Bénabou and Tirole (2003), Frey (1994, 1997), Francois (2000), Besley and Ghatak (2003), and related experimental economics findings.

### Implications for research and practice — policy recommendations and research agenda
- Research priorities:
  - Use actual expenditure and output/outcome data more extensively to assess allocative impacts and productive efficiency.
  - Design studies to distinguish policy-driven reallocations from endogenous changes (use of “constant policy” fiscal projection methods encouraged).
  - Focus on case studies of episodes of substantial short-run reallocation (e.g., Australia, New Zealand, United Kingdom) to disentangle causal impacts.
  - Investigate behavioral mechanisms linking performance budgeting to individual motivation and document both desired and perverse effects.
  - Assess the seriousness of jurisdictions’ commitments to performance budgeting rather than accepting claims at face value.
- Design guidance for practice:
  - Invest in performance measurement, management accounting, and improved performance information before or alongside implementation.
  - Combine performance budgeting with quality-regulatory frameworks and monitoring to mitigate perverse incentives.
  - Balance extrinsic incentives and intrinsic motivation: favor lower-powered, subjective, or team-based incentives where measurement is poor or goals are multiple.
  - Use performance measures even when imperfect; casemix experience suggests imperfect measures can still yield profound institutional effects if supported by complementary systems.
  - Tailor sectoral designs: casemix success in hospitals does not imply identical approaches for other sectors.

### Conclusions — concise findings
- The empirical literature does not demonstrate the systematic failure of performance budgeting.
- Where sufficient investment in performance measurement and supporting systems has occurred, performance information can be used in budgeting to improve allocative and productive efficiency.
- Government-wide empirical evidence is limited, methodologically constrained, and heavily based on opinion surveys; stronger, data-based, context-specific empirical research is needed.
- Sectoral casemix funding provides relatively strong evidence of productive efficiency gains with limited systematic adverse quality effects, contingent on not combining casemix with excessively severe immediate funding cuts and on complementary measurement and regulatory systems.
- Behavioral foundations and the interaction of intrinsic motivation with extrinsic incentives are central to effective design; the literature recommends carefully calibrated incentive systems adapted to measurement feasibility and mission-alignment.

*Source: Excerpted content from IMF working paper _wp05210 (selected sections and references as provided).*

### References..............................................................................................................

### _wp05210 - References

### I. Introduction: scope and motivation
- Performance budgeting has been an important theme of public expenditure management for decades, with a renewed wave of enthusiasm beginning in the 1990s and spreading from advanced to developing and transitional nations.
- The new initiatives are typically part of broader reforms changing public sector management and the public–private boundary, motivated by a perceived need to address disappointing public sector performance and to ensure fiscal sustainability—especially after fiscal consolidation episodes since the late 1980s.
- Noted skeptical assessments:
  - Schick (2003, p. 83): “efforts to budget on the basis of performance almost always fail.”
  - U.S. Congressional Budget Office (CBO, 1993, p. x): “little evidence of the much-touted advances in performance-based budgeting in local, state, and foreign governments. Performance measures have a limited ability to influence the allocation of resources.”
- Primary aims of the paper:
  - Review empirical literature on government-wide performance budgeting systems since 1990, distinguishing quantitative analyses (expenditure data, opinion surveys) from case studies.
  - Examine empirical literature on sectoral performance budgeting systems, with particular attention to the output-based “casemix” hospital funding model.
- Methodological caveats and challenges to causal inference:
  - Difficulty in obtaining conclusive empirical proof due to public sector performance measurement problems.
  - Three specific difficulties in assessing the impact of performance budgeting:
    1. Performance budgeting is part of broader management reforms, making attribution problematic.
    2. Success depends on contextual factors (political system, political culture, fiscal environment) that are hard to measure.
    3. The concept of performance budgeting has not always been well-defined.
- Research purpose: determine the extent to which negative views are validated by empirical literature, identify preconditions for effective performance budgeting, and suggest which forms work best.
- Paper structure (as presented):
  - Section II: definition and delineation of performance budgeting.
  - Section III: empirical literature on government-wide systems.
  - Section IV: sectoral systems and casemix hospital funding.
  - Section V: issues indicated by the empirical literature and suggestions for future research, including economists' roles.

### II. What is performance budgeting? — Definition and scope
- Working definition used in the paper:
  - Performance budgeting refers to procedures or mechanisms intended to strengthen links between the funds provided to public sector entities and their outcomes and/or outputs through the use of formal performance information in resource allocation decision-making.
  - “Formal” performance information includes performance measures, measures of the costs to particular parties of outputs and outcomes, and assessments of effectiveness and efficiency obtained through analytic tools.
  - Core objectives: enhanced allocative and productive efficiency in public expenditure.
- Relation to other definitions:
  - Close to OECD (“a form of budgeting that relates funds allocated to measurable results” (OECD, 2003b, p. 7)) and U.S. General Accounting Office (“the concept of linking performance information with the budget” (GAO, 1999a, p. 4)).
  - Encompasses classic U.S. models (Planning, Programming and Budgeting System—“program budgeting”; Zero-Base Budgeting; Hoover Commission proposals) and more recent international models (New Zealand and Australia output-purchase systems; U.K. Public Service Agreement—PSA).
- Terminology notes:
  - No distinction drawn between “performance budgeting” and terms like “performance-based budgeting,” “results-based budgeting,” or “performance funding.”
  - Outputs: goods/services provided to external parties; outcomes: desired changes induced by outputs and other factors; inputs: resources used to produce outputs.
  - Productive efficiency defined and decomposed into “technical efficiency” and absence of waste (X-inefficiency).
- Scope beyond central budget formulation:
  - Definition includes resource-allocation decisions within agencies and funding distribution to subordinate public entities.
  - Examples: ministry-level program budgeting marginal analysis, casemix hospital funding, output funding plus outcome bonuses for universities—termed “sectoral” funding systems.
  - Rationale: allocative efficiency arises at government-wide budget formulation, within agencies, and via service (output) design decisions; thus limiting empirical focus to central budget impacts is unduly narrow (Joyce, 2003, p. 7).
- Three levels of allocative decision-making (economist’s concept of allocative efficiency):
  - Resource allocation decisions embodied in the government-wide budget.
  - Resource allocation decisions within agencies.
  - Service (output) design decisions which affect program effectiveness.

### Key procedures and mechanisms used to link performance information to funding
- Classification of expenditure into “programs” with common objectives to facilitate prioritization and accountability (core idea of program budgeting; terminology varies).
- Setting of outcome or output targets linked to funding levels:
  - Example: U.K. PSA system emphasizing outcome targets and tougher performance targets as a condition for additional budget funding.
  - OECD (2003b, p. 9) survey: of 27 OECD nations responding, 12 asserted they set performance targets as part of the budget process, and in 9 of these the results were reported ex-post to the finance ministry.
- Formal agreements (“contracts”) that articulate funding-linked targets and sometimes consequences of delivery performance.
- Formula-based budget estimation based upon outputs: estimate expenditure by multiplying forecast/planned output volume by a per-unit funding amount (Hoover Commission approach included support services and activities as well).
  - Examples: school funding formulas based on student course-years; U.S. Working Capital Funds that base internal support services funding upon unit costs.
- Formularized output-based supplementary funding arrangements for unanticipated extra demand:
  - Example: Australia’s 1990s “workload agreements”/“purchasing agreements” where the immigration department received additional funding on a per-detainee/day basis for mandatory detention of illegal immigrants (DIMA/DOFA, 2001).
- “Output-purchase” budgeting:
  - Payment for outputs actually delivered.
  - Casemix funding: public hospitals funded principally on the services they deliver, with different prices set for every output type as defined by the “DRG” output classification system.

*Source: _wp05210 - References..............................................................................................................*

### Section V). This means that hospitals make a loss (profit) if the actual cost of the

### _wp05210 - Section V). This means that hospitals make a loss (profit) if the actual cost of the

### Performance budgeting mechanisms and examples
- Payment-by-results / casemix funding:
  - Hospitals are paid only for the services they deliver; "hospitals make a loss (profit) if the actual cost of the service delivered exceeds (is less than) the price they receive."
  - A strong financial incentive for efficiency is created assuming appropriately calculated prices.
  - New Zealand developed an “output purchase” performance budgeting system in the 1990s; Singapore adopted a similar “Budgeting for Results” system in the mid-1990s (Jones, 1998).
- Outcome/quality-based performance bonus funding:
  - Explicit funding component awarded based on achievement of outcomes and/or quality outputs.
  - Can be formula-based (e.g., university bonuses based on graduate employment rates or post-graduation salary levels) or judgmental via performance ratings.
  - Louisiana and a handful of other U.S. states have made provision for explicit performance bonus payments to state departments assessed as having performed well.
- Procedures to feed performance information into budget priority-setting:
  - Agencies may be required to justify budget bids using performance measures and other relevant performance information.
  - Finance ministries or central agencies may conduct program performance ratings or program evaluations for use in the budget process.
  - U.S. PART (Program Assessment Rating Tool) system: Office of Management and Budget (OMB) budget examiners rate in each annual budget round the performance of 20 percent of government programs; these ratings are expressly intended to inform resource allocation decisions in the president’s budget (Executive Office of the President/OMB 2002).
- Core institutional ingredient:
  - Reduction of centrally determined controls over the mix of inputs agencies employ (move away from "line item" controls toward allocations by outcome/output).
  - Consensus that tying managers’ hands on input mix guarantees technical inefficiency and high costs.

### Misconceptions about performance budgeting
- Narrow definitions:
  - Some define performance budgeting as merely reporting performance measures in budget documents; the paper’s conception focuses on mechanisms for using performance measures and other performance information in funding decisions.
  - Equating performance budgeting solely with the use of performance measures can lead to expectations of a direct or mechanical link between measures and budget decisions, which is often unreasonable.
- Past-performance vs. future-performance links:
  - Not all performance budgeting systems create ex post budgetary rewards or penalties based on past performance.
  - Some systems link funding to expected future performance (targets) and may use non-budgetary sanctions for failures.
  - Program budgeting traditionally is "foreign" to the idea of budgetary rewards for good past performance.
- Central planning critique:
  - Performance budgeting is sometimes misconceived as central planning of allocative choices. Contemporary recognition of information asymmetries leads central decision-makers to focus on policy/administration rather than operational detail.
  - Agency appropriations often take aggregated forms giving agencies freedom to shift money within and, with agreement, between programs.
- Distinction from broader managing-for-results reforms:
  - Performance budgeting is specifically concerned with the budgetary use of performance information (budget formulation and execution), and is only one component of managing-for-results.
  - Definitions that equate performance budgeting with broader strategic planning or GPRA-style reforms may not pertain directly to budgeting.

### Why performance budgeting? Objectives and motivations
- Core objectives:
  - Enhanced allocative efficiency and productive efficiency in public expenditure.
- Allocative efficiency motivations:
  - Address perceived insufficient responsiveness of expenditure allocations to changing social needs and priorities.
  - Counteract excessive budgetary "incrementalism" where base funding is routinely renewed without reappraisal.
  - Example: New Zealand reforms of the late 1980s and 1990s aimed to "assist the government to translate its strategy into action" more effectively (New Zealand Treasury, 1996, p. 7).
  - Improve responsiveness of resource allocation to short-run unexpected fluctuations in needs/demand in crucial areas.
- Performance pressure motivations:
  - Since 1990 most models view the budget process as a means to increase pressure on agencies to lift performance, using budget-linked targets, prices and performance agreements to strengthen ex ante linkage between results and funding.
- Productivity dividend and fiscal sustainability:
  - If performance budgeting boosts productive efficiency, aggregate public expenditure should, ceteris paribus, be smaller.
  - The "productivity dividend" can be used to keep the tax burden down or fund new service priorities.
  - Improved expenditure prioritization can make targeted spending cuts easier than across-the-board cuts (e.g., British efficiency targets in 1998 aimed "to ensure more resources go direct to front line services").

### Performance budgeting in the broader reform context
- Interaction with privatization and competition:
  - Performance budgeting reforms are one element in a wider response that includes privatization and increased private provision of publicly-funded services.
- Managing-for-results relationship:
  - Managing-for-results = use of formal performance information to improve public sector performance; emphasizes clarity about outcomes, links between outputs and outcomes, ex ante performance expectations, performance-based rewards and sanctions, and "let the managers manage."
  - Performance budgeting is a component of managing-for-results but distinct because of its budgetary focus.
- Synergies and dependencies:
  - Efficacy of performance budgeting likely depends upon complementary reforms (e.g., linkage between agency-level output/outcome targets and individual/work-group targets, strategic human resources management).
  - Nonbudgetary uses of performance information (strategic planning, nonfinancial rewards/sanctions) can be important complements.

### Key reasons performance budgeting might not work
- Political and incremental constraints on allocative efficiency:
  - Budgeting may be inherently political; "incrementalism" can dominate, limiting rational prioritization.
- Measurement and specification limitations:
  - Outcomes and their relation to outputs may be difficult or impossible to clearly specify in many public programs.
  - Performance measures are inherently imperfect; targets and price incentives linked to imperfect measures can produce adverse behavioral distortions (Smith, 1995).
  - Output measures often fail to capture quality or outcomes, risking erosion of quality if focus shifts narrowly to measurable outputs.
  - Quotation: "when the output [i.e., outcome or output] is difficult to measure, as is true in most government bureaucracies ...installation of specific goals may focus effort but may send the bureaucrats marching in the wrong direction" (Heckman, Heinrich, and Smith, 1997, p. 394).
- Uncertainty about appropriate budgetary incentives:
  - Debate persists about the role and appropriateness of financial rewards and sanctions for agency performance in government-wide systems.
  - Evidence is mixed and practices for applying incentives/disincentives to change performance have proven difficult (OPPAGA, 1997, p. 23).

*Source: _wp05210 - Section V). This means that hospitals make a loss (profit) if the actual cost of the service delivered exceeds (is less than) the price they receive.*

### references to the Soviet central planning experience, which abounds with delightfully

### _wp05210 - references to the Soviet central planning experience, which abounds with delightfully

### Critics, ethical concerns, and motivational ambiguity
- Critics compare target-based public management systems to Soviet central planning, citing anecdotes (e.g., the nail factory producing absurdly large nails) and labeling systems such as the British PSA as “neo-Stalinism” (Keaney, 2001).
- Concern that linking funding or financial incentives (agency or individual level) directly to imperfect performance measures can “crowd out” ethical/altruistic motivations deemed important for public sector performance.
- Performance budgeting models rarely specify how they are expected to affect individual work motivation:
  - The British National Audit Office identified a missing link in the PSA system concerning motivation (NAO, 2001, p. 34).
  - Even where agency-level funding “rewards” are emphasized, it is often unclear how these rewards are meant to motivate individual behavior.
  - If motivation is intended to operate through individual performance pay linked to performance budgeting, concerns about behavioral distortions and crowding out of “intrinsic” motivation become particularly compelling.
  - If motivation is not through “extrinsic” incentives, then the motivating force underpinning performance budgeting remains unspecified.
- The review emphasizes the need to assess both evidence of desired behavioral impacts and evidence of perverse and dysfunctional effects.

### Methodological challenges in assessing efficacy
- Assessing government-wide performance budgeting efficacy is methodologically formidable:
  - Allocative efficiency, program effectiveness, and productive efficiency are influenced by many variables; performance budgeting is only one.
  - Political changes, nonbudgetary management reforms (e.g., strategic human relations management), and institutional features can all affect outcomes and interact with performance budgeting.
  - Measurement problems are often severe; diversity of performance budgeting forms means evaluations should target specific variants rather than “performance budgeting” in general.
- Given complexity and limited experimental possibilities, theory and policy logic are important; nevertheless, empirical, “evidence-based” assessment remains crucial.
- The empirical evidence is considered under three headings: budgetary allocation; aggregate expenditure; and productive efficiency and program effectiveness.

### Empirical evidence on budgetary allocation impacts
- Key empirical questions:
  - Can performance budgeting improve allocation of public expenditure between competing purposes at functional and program levels?
  - Is performance budgeting more effective when combined with other budgeting or results-oriented reforms?
  - What pre-conditions or circumstances affect allocative efficacy?
- A threshold methodological obstacle: absence of a practicable way to measure degree of allocative (in)efficiency for a given budget allocation.
  - It is possible to measure changes in expenditure allocation, but there is almost no literature that starts from observed reallocations and assesses the contribution of performance budgeting to those reallocations.
  - The primary literature focus is subjective assessments of allocative impacts rather than objective attribution.
  - The budgetary incrementalism literature analyzes allocation changes but does not assess performance budgeting’s impact.

- Subjective assessments (surveys and case studies, largely U.S.-focused):
  - Four surveys asked whether performance measures led to changes in government-wide budget allocations/appropriations:
    - A 1999–2000 survey by Melkers, Willoughby, and others (on behalf of the GASB) found that only 7.5 percent of state budget office officials (executive and legislative) believed performance measures were “very effective” or “effective” in changing budget appropriations.
    - Jordan and Hackbart’s 1997 survey: officials representing 29 out of 45 respondent states agreed that “achievement of performance standards affects budget recommendations in the Governor’s Executive Budget,” and 33 out of 46 claimed that performance affects funding in the next fiscal year.
    - Poister and Streib (1999) found that 60 percent of city managers from jurisdictions with “centralized, citywide performance measurement systems that incorporate most departments and programs” believed performance measures brought about moderate or substantial changes in city budget allocations.
  - Eight surveys asked about perceptions of the use (rather than impact) of performance measures in budget/resource allocation decisions:
    - Four of these eight provide information on perceived use in government-wide budget allocation decisions.
    - The 1999–2000 Melkers, Willoughby et al. survey found only 26.3 percent of state executive budget offices considered output and outcome measures to be either “important” or “very important” in budget appropriation decisions.
    - Jordan and Hackbart reported that 23 of 41 respondent state executive budget offices believed performance indicators were an important tool in budget allocation decisions.
    - Poister and Streib reported that almost two-thirds of city managers from the described jurisdictions held a similar view.
    - Lee (1995) reported that 30 percent and 18 percent of state budget offices asserted that “substantial use” was made of output or outcome measures respectively in the formulation of the executive budget; Lee also reported some decline in the use of such information in budgeting from 1990 to 1995.
  - Surveys addressing internal resource allocation within agencies:
    - A GAO survey of federal government managers in 2000 found 43 percent of respondents reported using performance measures to a “great” or “very great” extent in making resource allocation decisions.
    - Poister and Streib reported almost two-thirds of respondents from jurisdictions with citywide systems believed performance measures were important or very important for budgeting purposes.
    - A 1996 GASB survey (Melkers and Willoughby) reported that among respondents from state departments claiming to have developed performance measures:
      - 44.7 percent claimed output measures were used for resource allocation purposes.
      - 41.7 percent claimed outcome measures were used for resource allocation purposes.
    - For municipalities in that survey:
      - 20.1 percent and 17.7 percent reported use of output and outcome measures respectively for resource allocation.
    - For counties, the corresponding figures reported were 33.6 percent (output) and 33.6 percent (outcome) as stated in the same passage.

### Implications highlighted for further research and policy
- Empirical literature is limited and methodologically constrained; more rigorous approaches are needed to:
  - Measure allocative efficiency changes and attribute those changes to performance budgeting.
  - Distinguish effects of different forms of performance budgeting.
  - Investigate behavioral mechanisms linking performance budgeting to individual motivation and performance, and document perverse/dysfunctional effects as well as desired impacts.
- Policy logic and theory remain important supplements to empirical work given causal complexity, but evidence-based policy requires improved empirical methods and data focused on specific forms and contexts of performance budgeting.

*Source: Excerpt from the specified PDF content unit.*

### 28.1 percent. Results of earlier surveys by the GAO in 1996–97 and by Wang in 1996 were

### _wp05210 - 28.1 percent. Results of earlier surveys by the GAO in 1996–97 and by Wang in 1996 were

### Survey results on internal resource allocation and program budgeting
- Surveys pertaining to internal resource allocation uses of performance information appear possibly to be more positive than those relating to resource allocation at the level of the government-wide budget, although the data does not permit strong conclusions.
- Very little survey or case study work of this type exists outside the United States.
- Kluvers (2001, p. 40): of respondent local government entities in Victoria, Australia, with program budgeting systems, approximately half believed that it influenced the allocation of resources.
- Bellamy and Kluvers (1995, p. 52): earlier survey of the same group reported similar results.
- Poister and Streib survey: directed questions on resource allocation uses and effects to officials from jurisdictions claiming to have “centralized, citywide performance measurement systems that incorporate most departments and programs”; this survey reports the most positive views.

### Interpretations of survey evidence on the efficacy of performance budgeting
- Jordan and Hackbart (1999, p. 85) suggest: “performance budgeting may impact the appearance and preparation of the budget document, but the outcome in terms of funding is not significantly (if at all) affected.”
- The paper notes this conclusion is not clearly justified by those surveys; negative perceptions may reflect the early state of development or absence of performance measurement and performance budgeting systems in many jurisdictions.
- Development of good performance measures—including output and outcome measures—is a necessary precondition for performance budgeting and takes time and serious effort.
- Melkers and Willoughby (1998) reported adoption by the end of the 1990s of GPRA-style performance measurement and strategic planning systems in most U.S. states, but their definition of “performance budgeting” is problematic because it does not pertain to the budget formulation process.

### GASB case studies and self-reporting concerns
- Seventeen standardized GASB case studies examined use and effects of performance measures in U.S. states, cities, and counties.
- In only 6 out of the 17 jurisdictions did it appear that performance measures were significantly used in, and impacted upon, resource allocation decisions.
- A majority of the 17 jurisdictions had apparently only recently initiated moves to use measures in the budget process and most started from a weak performance measurement base.
- Of eight jurisdictions that were not newcomers to performance budgeting (introduced no later than the mid-1990s), six claimed substantial effect on budget allocations (those six: Oregon, Texas, Iowa, Winston-Salem, Prince William County and Sunnyvale).
- The GASB case studies reduce but do not eliminate self-reporting bias; CBO (1993, p. xi) cautioned that self-reported surveys are limited in providing detailed and verifiable information in complex areas such as performance measurement.

### Case study evidence on evolving roles of performance measures
- Oregon case study: initial attempt to establish mechanical linkages between outcomes and funding failed, but “the use of performance measures in the budget process seems to have evolved into one of providing elected officials with information that will assist them in decision making and in understanding the role of various state programs in achieving the results deemed important by the legislature and the governor, not in directly helping establish funding levels” (Fountain, 2000, pp. 11–12).
- Texas case study: “It is difficult to link actual dollars allocated to specific performance measures and convincingly argue that X dollars were appropriated because of Y performance. However, there is evidence that budget decisions include discussions about agency and program performance” (Tucker, 2000c, p. 2).

### Quantitative evidence on expenditure allocation impacts
- Only one paper identified that quantitatively examines impact of performance budgeting on expenditure allocations: Reddick (2003a).
  - Reddick tests impact of “rational budgeting” (program budgeting, zero base budgeting, or biennial budgeting) on expenditure levels in eight broad functional sectors.
  - Finding: rational budgeting associated with restraint of spending levels in only a minority of functional spending sectors.
  - Caveat: Reddick tests levels of expenditure rather than reallocations between functional areas; interpretation depends on whether one expects performance budgeting to restrain spending within each sector.
- Brumby, Edmonds and Honeyfield (1996) on New Zealand:
  - After introduction of “financial management reform” (performance budgeting plus broader reforms) in the late 1980s, there were “bold reprioritization decisions” consistent with clearly described government budgetary preferences.
  - The paper concludes there is “insufficient evidence, at this stage, to form a judgment concerning financial management reform’s contribution to improved prioritization.”

### Budgetary incrementalism and allocative flexibility
- Budgetary incrementalism: budgeting characterized by “inattentiveness to the (budgetary) base”; present allocation strongly predicted by previous year’s expenditure/base.
- Empirical literature:
  - Pre-1980s: mixed results on incrementalism.
  - 1980s: most studies concluded the U.S. federal budgetary process is no longer incremental, showing significant periodic shifts in budgetary expenditure allocations (Berry, 1990, p. 168).
- Post-1990 studies:
  - Dezhbakhsh, Tohamy and Aranson (2003): analyze U.S. federal nondefense expenditure 1946–1996, firmly rejecting incrementalism and identifying triggers for major shifts.
  - Reddick (2003b): U.S. federal expenditure 1968–1999 found most broad functional categories follow extrapolation of past spending decisions.
  - Reddick (2002a): comparison of U.S., U.K., and Canada (1950–2000; Canada 1961–2000) found no overwhelming support for incrementalism, but greater incrementalism in the United States.
  - Boyne, Ashworth and Powell (2000): study of 403 English local governments (1982–1996) casts considerable doubt on incrementalism.
  - Reddick (2003a): U.S. state budgets 1960–96 concluded incrementalism had “relatively low explanatory power” for broad functional categories; Canadian provincial studies found some degree of incrementalism.
- Conclusion: literature does not support belief in a strong form of incrementalism; some degree of incrementalism is not incompatible with effective performance budgeting and may coexist with “punctuated” incrementalism.

### Methodological issues in incrementalism and measurement
- Challenge of distinguishing endogenous from exogenous changes in expenditure allocations (endogenous = shifts that would occur when expenditure policies remain stable).
- High-level aggregation in incrementalism literature (SNA broad categories) may miss reallocative flexibility at less aggregated levels, including internal resource allocation within agencies.
- Recent developments in medium and longer-term fiscal projections that quantify endogenous (“constant policy”) fiscal impacts are relevant to future research.

### Institutional and political factors affecting allocative efficacy
- Developed performance measurement system is important for success of performance budgeting.
- Distribution of budgetary powers matters:
  - Parliamentary systems concentrate budgetary power in the executive/Cabinet; if party discipline is strong, executive budgets often receive parliamentary endorsement with little amendment.
  - U.S. federal and state systems allocate substantial budgetary power to legislatures; the executive budget may be significantly changed by legislative appropriations committees.
  - Some U.S. states have no executive budget; agencies present bids directly to legislature.
- Implications:
  - In parliamentary systems, Cabinet can impose a single set of priorities; in U.S. system, neither legislature nor president/governor can reliably do so.
  - Performance budgeting conducted solely within the executive is likely to be disregarded in legislative decision-making—seen as a key reason why pre-GPRA reforms failed.
- GPRA aims to build consensus between legislative and executive arms via participatory processes and joint strategic planning, but achieving such consensus is extraordinarily difficult.
- CBO (1993, pp. 11–12): performance management systems seem to work best in council/manager local governments and national governments with parliamentary systems, where concentration of political power in one branch encourages goal development and meaningful measures.

### Evidence on locus of resource allocation effects
- U.S. survey data impression: performance information may have more impact on internal agency resource allocation than on government-wide budget allocations.
- Hatry, Morely, Rossman and Wholey (2003): selection of U.S. federal program case studies supports this impression.
- Joyce (2003): suggests researchers seeking resource allocation effects of performance budgeting at the government-wide level in the U.S. may be looking in the wrong place and should focus on internal resource allocation impacts.

### Aggregate expenditure impacts
- Little literature on performance budgeting’s impact on aggregate spending.
- Reddick (2003a): testing impact on aggregate expenditure, concludes “rational [budgeting] reforms have been successful in reducing total expenditures” (2003a, p. 336), i.e., states with program budgeting, ZBB or biennial budgeting more likely to have restrained aggregate expenditure after controls.
  - Caveat: absence of causality testing leaves open alternative explanations (e.g., governments under fiscal pressure adopt performance budgeting).
- Brumby, Edmonds and Honeyfield (1996): New Zealand post-reform succeeded in reducing central government expenditure as a proportion of GDP; authors say evidence is consistent with, but does not conclusively establish, financial management reform having made it easier to control public expenditure.

### Impacts on productive efficiency and program effectiveness
- Almost no literature measuring improvements in output costs and outcomes attributable to performance budgeting.
- Brumby, Edmonds, and Honeyfield (1996): unit cost data for four “process” outputs from three ministries showed data for three of four outputs suggesting significant productivity gains.
  - Authors caution limited sample size and that gains may be larger for more measurable outputs because “what is measured is more likely to be managed.”
  - They cannot attribute gains specifically to performance budgeting versus the broader package of reforms or other factors (ministerial determination, technological change).
- Measurement problems:
  - Need for good “before and after” data; often robust performance measures did not exist prior to implementation, so “before” data is frequently missing.
  - Dominance of subjective-assessment literature (surveys) rather than objective measures of efficiency/effectiveness.
- The U.S. survey literature focuses on use and effects of performance measures, not directly on impacts of performance budgeting on efficiency and program effectiveness.

*Source: Excerpt from the supplied IMF working paper content unit.*

### 46.4 percent of those from cities with citywide performance measurement systems believed

### _wp05210 - 46.4 percent of those from cities with citywide performance measurement systems believed

### Impacts of performance measurement and managing-for-results (survey evidence)
- 46.4 percent of those from cities with citywide performance measurement systems believed that performance measures had a moderate or substantial impact in reducing the costs of services.
- 71.6 percent of those from cities with citywide performance measurement systems believed that performance measures had a moderate or substantial impact on the quality of services.
- The 1999–2000 GASB survey concluded that “more than half of all respondents agree that the implementation of performance measures has increased both the efficiency and effectiveness of their various governmental programs” (Melkers et al., 2002, p. 32).
- GAO 1996–97 survey of federal managers’ view on GPRA: 42 percent of respondents indicated that they believed that GPRA had improved programs to a “moderate or greater extent” (GAO, 1997b, p. 86).
  - The GAO survey sampled managers and supervisors in 24 agencies, obtaining useable responses from about 72 percent of the 1,300 persons surveyed; 41 percent of these managers expressed an opinion on the extent to which their agencies’ efforts to implement GPRA to date had improved their agencies’ programs, and it was 42 percent of these who responded as reported above.
- 1997 survey of U.S. state budgeters (Willoughby and Melkers, 2000): respondents rated effectiveness of a GPRA-type package on a four point scale from 1 (not effective at all) to 4 (very effective); mean ratings were:
  - 2.17 for “improving effectiveness of agency programs,”
  - 1.75 for “affecting cost savings,”
  - 1.79 for “reducing duplicative services.”
- Overall assessment: survey results suggest a significant impact of performance measures and managing-for-results strategies generally on effectiveness and program efficiency, but survey data are subject to self-reporting bias.

### Gaps in the literature on productive efficiency and program effectiveness impacts
- Identified empirical gaps:
  - Lack of empirical literature testing the importance of managerial flexibility in the use of inputs for realizing intended aims of performance budgeting.
  - Little evidence beyond anecdotes on unintended and perverse effects (e.g., quality- and outcome-eroding effects) of output-focused government-wide performance budgeting systems.
- Measurement difficulties contribute to the gap; survey and case study literature has not adequately addressed unintended or perverse effects.
- Stronger evidence exists for sectoral performance budgeting systems; the subsequent section examines one such system in detail.

### Sectoral performance budgeting systems: Casemix funding (overview and rationale)
- Casemix funding (DRG-based or similar output classifications) is one of the most long-established and well-researched sectoral performance-based funding approaches.
- DRG (diagnostic related groups) classify acute in-patient hospital treatments into categories “relatively homogenous in respect of the resources used” (Palmer and Reid, 2001), often called “iso-cost” output classes.
- Historical development:
  - DRGs developed in the 1970s for cost information and performance management.
  - U.S. federal government introduced Prospective Payment Systems (PPS) for Medicare in 1984, paying fixed prices per DRG output type.
- Effects of DRG-based pricing:
  - Any difference between actual cost and DRG price represents a loss or profit to the hospital.
  - Price-setting approach matters: setting prices near average penalizes (rewards) relatively inefficient (efficient) hospitals; setting prices near costs of most efficient hospitals penalizes even average inefficiency.
- International adoption: casemix funding adopted in Portugal (1990), Australia (from 1993), Norway (1997), Singapore (1997), and the United Kingdom (2004).
- Distinction in public hospital contexts:
  - Some public systems use fixed annual “global budgets” with DRG output measures determining budgets, creating financial penalties and rewards tied to delivering outputs below prevailing DRG “price.”
  - Prior to casemix, public hospitals commonly received annual line-item budget allocations with little link between funding and performance.
- Policy intent: introduce financial incentives to promote efficiency, increase treatment volumes, reduce waiting times, and free hospitals from bureaucratic input controls to promote managerial autonomy.

### Concerns and behavioral responses under casemix funding
- Major concerns at introduction:
  - Potential quality and outcome erosion because casemix payments focus on output quantity without direct payment for quality or outcomes.
- Potential dysfunctional behaviors induced by DRG pricing heterogeneity:
  - “Skimping” — under-provision of services to relatively high-cost patients within a DRG category.
  - “Dumping” — avoidance of patients whose expected costs exceed the DRG price (explicit or covert non-admission or inappropriate transfers).
  - “Creaming” — over-servicing or encouraging admission of relatively low-cost, profitable patients (part of “cream-skimming”).
- Mitigation measures implemented:
  - Quality-oriented regulatory structures and monitoring regimes to detect inappropriate conduct.
  - Input-based payments for anomalous “outlier” patients with costs greatly exceeding DRG price.
  - Progressive refinement of DRG classifications to reduce heterogeneity (e.g., disaggregating hip fracture DRG into multiple more “iso-cost” groups).
- Even with safeguards, casemix systems faced inherently imperfect performance measures; quality/outcome risks remained a central empirical concern.

### Empirical evidence on casemix funding impacts
- Coulam and Gaumer (1991) literature synthesis (up to 1991) findings:
  - The U.S. PPS had been highly successful in improving hospital efficiency and reducing rate of increase of Medicare health costs.
  - “None of the worst fears raised at the outset have been borne out by experience.”
  - No evidence of significant deterioration of health outcomes to that time.
  - No evidence of other feared perverse effects specific to the U.S. health system.
  - Evidence did not show a “large and systematic reduction in the rates of adoption of new technology.”
- Supporting qualitative evidence:
  - Survey of 10 pairs of “winner” and “loser” hospitals under PPS in the 1980s found many “loser” hospitals had excessive staffing that could be reduced “without sacrificing quality of patient care,” while “winner” hospitals reported lean staffing perceived to be about right (Bray et al., 1996).
- Literature coverage since 1991:
  - Relatively few empirical studies relevant to the period since Coulam and Gaumer’s review; much literature surveyed in later reviews pre-dates 1991.
- Important empirical caveat:
  - Some evidence suggested an increase in the degree of health status instability of patients at discharge after PPS introduction, raising concerns about discharge planning and reallocation of services to outpatient, home care, nursing homes, and other settings.
  - Reallocation of services driven by PPS may represent a re-specialization that contributed materially to efficiency improvements by shifting care to more cost-effective settings.

*Source: _wp05210 - 46.4 percent of those from cities with citywide performance measurement systems believed (IMF working paper content).*

### introduction of PPS. Over time, with the accretion of other confounding factors, this

### _wp05210 - introduction of PPS. Over time, with the accretion of other confounding factors, this

### Evidence on PPS / casemix funding impacts: cost containment and quality
- Empirical comparisons:
  - Boccuti and Moon (2003) compared Medicare cost increases with private insurers and concluded: “Medicare has proved to be more successful than private insurance in controlling the growth rate of health care spending per enrollee.”
  - Only one post-1991 US paper on quality/outcome effects of PPS identified: Meltzer, Chung, and Basu (2002) find some evidence of “skimping” in a study of California patients.62
  - Predominant contemporary view: “overall ...the system has worked pretty well in establishing better incentives without having a systemic adverse effect on quality.”63
- International and subnational studies:
  - Portugal: Dismuke and Guimaraes (2002) studied mortality rates for the most frequent non-obstetric DRG category before and after casemix funding and concluded “no evidence was found that the case-mix based payment system ...has had a pernicious effect on hospital quality as measured by hospital mortality.”
  - Portugal (technology use): Dismuke and Sena (1999) found casemix funding had a positive effect on productivity in the use of three of the most common diagnostic technologies.
  - Victoria, Australia (first Australian state to adopt casemix in 1993):
    - Young and Harris (1999) presented results “quite consistent with the expectation that casemix funding is providing hospital with appropriate incentives for cost efficiency,” while acknowledging methodological limitations.65
    - Victorian Auditor-General (1998) concluded the goal of “substantial efficiency gains ...has been effectively met” and pointed to official cost data showing dramatic cost reductions and changes in treatment patterns suggesting significant efficiency increases (1998, pp. 187–222). The Auditor-General also considered there had been an “overall decline in the quality of care” (1998, pp. 5, 65), a conclusion based on opinion surveys of hospital administrators and medical staff and therefore methodologically questionable.
    - One empirical study relevant to quality in Victoria: Brown and Lumley (1998) — patient satisfaction surveys of women who had just given birth — found a doubling of critical comment on quality of care; the 1993 survey was conducted 3 months after casemix introduction and the authors concluded “standards of care are being compromised in the current economic and policy environment.”66

### Explanations for the generally benign quality outcomes despite financial incentives
- Role of regulatory and quality-management structures:
  - Enhanced quality-oriented regulatory structures accompanied casemix funding and were supported by significant improvements in quality-related performance information and management processes (quality assurance, accreditation, public release of outcomes in a few instances).
  - These mechanisms illustrate synergy between performance budgeting and broader performance management frameworks.
- Role of socialized, value-based behavioral drivers:
  - Professional ethics and commitments to quality care act as countervailing forces to financial incentives that could degrade quality.
  - Coulam and Gamuer (1991, p. 65): “PPS in effect viewed hospitals, physicians, and others as buffers between the purely financial incentives of PPS and the patient needs for quality care.”
  - No empirical literature quantifying the contribution of altruistic, value-based drivers to safeguarding quality after casemix introduction, but the impression is that they “may have played a major part in the benign impact of the casemix funding system.”
- Interaction with funding levels and timing:
  - Victoria’s casemix introduction coincided with substantial immediate cuts to overall hospital funding driven by state-wide fiscal consolidation, making it difficult to disentangle casemix effects from funding-cut effects.
  - In the United States and many other countries, casemix/PPS was introduced designed to be initially cost neutral with savings realized over time as efficiency improved.

### Performance information and measurement lessons
- Strong empirical grounds suggest casemix funding delivered strong efficiency gains without demonstrable adverse effects on quality and outcomes, with qualifications.
- Twin pillars of success:
  - Performance information.
  - Clear procedures and mechanisms for use of that performance information in funding and managerial decisions.
- Key observations about measurement:
  - Introducing performance management and budgeting without large investment in improved performance information invites disappointment.
  - Casemix experience suggests it is not necessary to wait for near-perfect performance measures before putting systems to use. Coffey (1999): “even with seemingly inadequate databases, Medicare’s casemix reimbursement ...had profound effects.”
  - Hospital outputs are heterogeneous and not standardized mass-produced “widgets,” yet casemix succeeded despite measurement difficulties.

### Implications for sectoral performance budgeting more generally
- Caution on generalization:
  - Casemix funding is one example; efficacy and appropriate models of performance budgeting will vary across sectors.
  - What works for hospitals may not work for education, police, or other health subsectors.
- Perverse effects and sectoral variation:
  - Absence of pervasive perverse effects in casemix experience is notable but not necessarily generalizable.
  - Empirical evidence of perverse effects exists in other public sectors (school education; labor market programs such as JTPA) summarized in Propper and Wilson (2003) and Courty and Marschke (2003).
  - Historical record of target-setting perverse effects in Soviet-style centrally planned economies (Nove, 1984) reinforces risks.
- Role of altruism and intrinsic motivation:
  - Casemix experience suggests perverse effects may be substantially mediated by altruistic commitment of public sector employees to service quality.
  - Strength of altruistic commitment likely varies by sector and is particularly strong in hospitals due to professional socialization and severe ramifications of lapses in care.
  - Managing-for-results systems may reshape and balance altruistic commitment: practitioners may be myopically focused on individual clients; performance budgeting can communicate financial implications of clinical decisions and promote broader cost-effectiveness (Averill et al., 1996).
  - Evidence from JTPA labor market program:
    - Caseworkers manifest “strong preferences for serving the most disadvantaged (and least employable)” which counteracts “cream-skimming” incentives (Heckman et al., 1996).
    - However, this bias can misallocate resources toward clients least likely to succeed; performance-based funding acts as a partial check on such preferences (Heckman, Heinrich, and Smith, 1997, p. 393).
    - Conclusion: altruistic motivation and managing-for-results systems can be synergistic, moderating excesses of each other.

### Conclusions and overview of findings
- Main aim: assess what empirical literature shows about efficacy of performance budgeting and whether linking funding to results in government budgeting has failed.
- Empirical literature on government-wide performance budgeting is “disappointingly limited in scope and methodology” and does not provide a basis for strong general conclusions.
- With qualifications, casemix funding provides a robust example of sectoral performance budgeting delivering cost containment and efficiency gains without clear systematic adverse quality or outcome effects, especially when not combined with immediate severe funding cuts.
- Important factors in success include:
  - Investment in performance information.
  - Quality-regulatory frameworks and management processes.
  - Interaction of financial incentives with altruistic, value-based professional motivations.
  - Use of performance measures even when imperfect can still yield profound institutional effects.

*Source: _wp05210 - introduction of PPS. Over time, with the accretion of other confounding factors, this*

### conclusions about the efficacy of these systems. Nevertheless, it does appear to provide some

### _wp05210 - conclusions about the efficacy of these systems. Nevertheless, it does appear to provide some

### Empirical findings on performance budgeting
- The empirical literature does not demonstrate the failure of performance budgeting.
- Where necessary investment in the development of performance measurement and other performance information has been made, it is possible to use that information in budgeting to improve both allocative and productive efficiency.
- Much of the literature focuses on performance measures rather than performance budgeting per se; performance budgeting concerns the budgetary use of performance information generally, not only measures.
- Simplistic notions of the link between measures and budgets are not uncommon and have affected the empirical literature.

### Sectoral evidence: casemix funding of hospitals
- There is quite strong evidence of the efficacy of the sectoral system discussed—“casemix” funding of hospitals.
- Casemix funding has delivered strong productive efficiency gains and appears to have done so with surprisingly little in the way of perverse and unintended effects.69
- The casemix example illustrates the crucial role of:
  - performance measurement,
  - management accounting, and
  - the relaxation of input controls
  as underpinnings of successful performance-based funding.
- Caveat: casemix outcomes are “at least when not accompanied by excessively severe immediate cuts in hospital funding levels.”69

### Cross-national and methodological observations
- There is a dearth of empirical studies of non-U. S. government-wide systems.
- Paradox: the United States has historically led on performance budgeting but its system may be relatively unfavorable to the use of performance information in allocative decisions in the annual budget.
- Nations with concentrated budgetary power in the executive (Australia, New Zealand, United Kingdom) have undertaken major policy-driven budgetary expenditure reallocations with explicit procedures for linking funding and results; little close empirical study exists on the role of performance information in these reallocations.
- The literature has divergent and occasionally unclear definitions of performance budgeting; some claims of having introduced performance budgeting are accepted too readily.
- Strikingly little use has been made of actual expenditure data to assess allocative impacts, and of output/outcome measures to assess productive efficiency and program effectiveness; opinion surveys have been the primary research tool.
- Methodological and data-availability problems explain limited data use, but there is scope for more data-based empirical research.
- A threshold challenge is distinguishing policy-driven from endogenous changes in expenditure allocations; recent advances in “constant policy” fiscal forecasting and projection methods are potentially useful.
- The challenge of distinguishing policy-driven from endogenous changes is greater for medium and longer-term changes because many endogenous changes occur gradually; in the short run the problem is less severe and relates mainly to cyclical components of expenditure.
- Case studies of episodes of substantial short-run budgetary expenditure reallocation (e.g., Australia, New Zealand, United Kingdom) may be particularly useful.
- Disentangling the impact of performance budgeting from other factors (such as changes in political configurations) is challenging; statistical methods may have limited help, and case studies of concrete examples may be more fruitful.
- For impacts on productive efficiency and program effectiveness, it is probably not practicable to distinguish the impact of performance budgeting from that of managing-for-results reforms more generally.
- As quality of performance measures improves (including outcomes, quality, client satisfaction surveys), empirical analysis of managing-for-results impacts and tests for perverse/unintended effects from target-setting will become increasingly possible.

### Behavioral foundations: motivation and incentives
- A weak link in the literature is the connection with behavioral theory; there is a need for explicit articulation of how performance budgeting systems are assumed to motivate public officials.
- Key question: Are motivating forces assumed to arise from links between performance budgeting and stronger results-based extrinsic incentives (e.g., performance pay)? If not, what is the mechanism?
- Effective systems will reflect an accurate appreciation of what motivates people in organizational contexts, particularly in the public sector.
- Recent decades saw a shift toward viewing human organizational behavior as fundamentally self-interested (homo economicus), replacing earlier assumptions of “public spirited altruists.” This influenced managing-for-results and performance budgeting, including the view that funding “rewards” for agencies is appropriate.
- Organizational economics warns of perverse and dysfunctional responses to high-powered extrinsic incentives under poor measurability, uncertainty, and multiple goals (“multi-tasking”).
  - The literature concludes greater reliance generally should be placed upon lower-powered, “fuzzier” extrinsic incentives (career systems, subjective evaluation, team-based bonuses) or fixed remuneration combined with intensive monitoring (Gibbons, 1998, p. 120).
  - These problems tend to be particularly severe in the public sector; additional public-sector-specific problems include the presence of multiple principles.
  - From this literature it follows that the optimal role for high-powered extrinsic incentives is generally less in the public sector than in the private sector (Tirole, 1994; Dewatripont, Jewitt, and Tirole, 1999; Dixit, 2002; Grout and Stevens, 2003).
- The pure self-interest postulate is incorrect in the public sector; “public service motivation” (internalized commitment to a mission) and intrinsic motivation matter.
  - Empirical literature provides substantial evidence that intrinsic motivation is significantly more important in the public sector: government employees are on average significantly more driven by a sense of mission than private sector employees, and among public sector workers “service-oriented” employees are “more productive than economic-oriented employees” (Crewson, 1997, p. 515).
- Excessive reliance upon extrinsic incentives can “crowd out” intrinsic motivators, reducing performance (e.g., Osterloh and Frey, 2000). The greater the importance of intrinsic motivators, the greater the risk of crowding out.
- Economists have recently begun addressing intrinsic motivators and “public service motivation”:
  - Bruno Frey introduced intrinsic/extrinsic motivational theory to economics and coined “crowding out” terminology in this context.
  - Bénabou and Tirole (2003) develop economic theory on the interplay of extrinsic incentives and intrinsic motivators and circumstances for crowding out.
  - Francois (2000) models how business-derived management/incentive practices can diminish effort based on public service motivation.
  - Besley and Ghatak (2003) model mission-commitment matching in recruitment and the role of lower reliance on extrinsic motivators.
  - Experimental economics reinforces risks that high-powered monetary incentives can produce worse outcomes (see research summarized in Bénabou and Tirole (2003, p.490)).
- Nonetheless, this literature does not argue against any increase in extrinsic motivators or greater use of performance measurement.
  - Public officials are driven by a mix of motivations; the optimal system mixes extrinsic incentives and public service motivation.
  - The traditional low reliance upon extrinsic incentives in public administration may not have been optimal; some empirical evidence supports too little past reliance (Burgess and Ratto, 2003, p. 298).
  - There is no basis to assume expansion of extrinsic incentives will necessarily come at the expense of intrinsic motivation.
  - Evidence (Deckop, Mangel and Cirka, 1999) suggests in organizations with high value-alignment the right performance pay system can complement intrinsic motivation (“organizational citizenship behavior”).
  - Taylor and Pierce (1999) provide some evidence of a similar effect in a public sector context.
- Presence of strong intrinsic motivations may permit a greater role for the right types of extrinsic performance incentives even when measurement problems, uncertainty, multi-tasking, and multiple principles exist, because mission commitment can partly neutralize exploitation of imperfect measures.
- The casemix funding example demonstrates that anticipated dysfunctional/perverse responses to financial incentives linked to imperfect measures were not realized in that context.

### Implications for research and practice
- Future researchers should assess how serious a jurisdiction’s commitment to performance budgeting actually has been rather than accepting claims at face value.
- More use should be made of actual expenditure and output/outcome data where possible; data-based empirical research has scope to expand despite methodological challenges.
- Study of substantial short-run expenditure reallocations (e.g., Australia, New Zealand, United Kingdom) may yield particularly useful insights.
- Case studies of concrete examples of performance information use in major expenditure reallocation decisions are likely to be more fruitful than purely statistical approaches for disentangling causal impacts.
- Continued improvement in measurement of outcomes, quality, and use of client satisfaction surveys will enable stronger empirical tests of managing-for-results and of perverse/unintended effects from target-setting.
- Design of performance budgeting systems should balance extrinsic incentives and intrinsic motivation, favor lower-powered/fuzzier incentives or subjective/team-based measures where measurement is poor or goals are multiple, and pay attention to organizational commitment and mission alignment.

*Source: _wp05210 - conclusions about the efficacy of these systems. Nevertheless, it does appear to provide some*

### Section IV provides, perhaps, some support for this proposition—prior predictions of serious

### _wp05210 - Section IV provides, perhaps, some support for this proposition—prior predictions of serious

### Imperfect output measures and observed effects
- Prior predictions of serious perverse and dysfunctional effects from the highly imperfect output measures upon which the system was based appear not to have been realized.
- Possible explanation: commitment of health practitioners to quality of care and professional standards.
- Casemix funding system may be a particularly positive example of this neutralization effect, given some empirical evidence of “gaming” responses to the choice of performance measures used in other sectoral performance budgeting and management.
- Caveat: presence of some undesirable responses to imperfect measures does not imply there has not been a net improvement in performance from the use of performance information in management and budgeting.
- Some supposedly perverse effects may instead reflect legitimate policy choices (for example, a desire to weight efficiency considerations more vis-à-vis equity).

### Public service motivation, mission-orientation, and goal alignment
- Building and molding public service motivation and mission-orientation are central to effective public management.
- Managing-for-results and performance budgeting systems should be evaluated explicitly for their impact on public service motivation and mission orientation.
- It is necessary both to nurture/reinforce mission-orientation and to be able to change and re-mold it because public officials’ conception of their mission may diverge from political masters or the broader community for reasons including:
  - ideological orientation (and the failure to adapt that orientation to changing circumstances),
  - professional values/norm,
  - myopia (e.g., failure to see trade-offs such as equity versus efficiency in the context of limited resources).
- Managing-for-results processes—making desired outcomes explicit and linking outputs and activities to those outcomes—help improve “goal alignment.”
- Target-setting, preferably closely linked to budget formulation, plays an important signaling role.
- Example: Courty and Marschke (2003b) present a model where imperfect performance measures communicate priorities between competing objectives to agents in a multi-tasking environment, even absent extrinsic incentives.
- Measures and targets that capture organizational success in achieving outcomes can focus attention on the bigger picture and partially offset myopia among people operating as small parts of larger production processes.

### Managerial freedom, measurement limits, and the role of intrinsic motivation
- The managerial freedom theme of managing-for-results advocates relaxation of input controls and relies on accountability for results (outcomes and outputs) to replace process controls.
- Limitation: serious constraints on our ability to measure results in many parts of the public sector.
- Organizational economics proposition: where results delivered by self-interested agents are hard to measure, intensive monitoring and control of agents’ effort (processes) is likely if shirking is to be limited—this undercuts the managerial freedom argument.
- Counterpoint: strong public service motivation and mission-orientation (altruistic motivators) increase the likelihood that de-control of processes and inputs will improve performance rather than increase shirking.

### Need for better theory and empirical evidence
- Analysis is acknowledged as speculative; Wright (2001) notes public sector work motivation has received far too little scholarly attention.
- Call for: much better theory with stronger empirical foundations to improve understanding of civil service motivation and behavior (Crewson (1997, p. 500) quoted).
- Worldwide experiments changing motivators and incentives faced by public officials are underway and will provide data to improve design of managing-for-results systems and performance budgeting.

### Appendix I — U.S. studies relevant to the efficacy of performance budgeting: selected empirical findings
- Melkers, Willoughby et al., Date when empirical work carried out: 1999–2000; Level: State and Local.
  - Nature of study: Survey of (i) “state budget offices” (responses from 36 states); (ii) state agency officials with budgeting and/or performance measurement responsibilities (responses from 48 states); (iii) budget officers & department heads of city and county governments (responses from 253).
  - Findings: The percentages of respondents who asserted that performance measures were “very effective” or “effective” in changing budget appropriations were 7.5 percent (state budget offices), 24.2 percent (state agencies) and 16.9 percent (city/county) (Melkers, Willoughby, James, Fountain and Campbell, 2002, p. 21).
- Melkers and Willoughby, Date: 1997; Level: State.
  - Nature of study: Survey of executive budget offices and legislative budget officials. 104 responses from 49 states.
  - Findings: In response to the proposition that “some changes in appropriations are directly attributable to outcomes from the implementation of performance-based budgeting in my state,” 9.4 percent of respondents strongly agreed and 29.7 percent agreed. The figures were significantly lower for legislative branch respondents than for those from the executive branch (Melkers and Willoughby, 2001, p. 61). When asked to rate the effectiveness of performance-based budgeting in “changing appropriation levels” on a four point scale from 1 (not effective at all) to 4 (very effective), the average rating was 1.54 (Willoughby and Melkers, 2000, p. 114).
- Jordan & Hackbart, Date: 1997; Level: State.
  - Nature of study: Survey of executive budget offices. 46 state responses.
  - Findings: 29 out of 45 respondent states agreed that “achievement of performance standards affects budget recommendations in the Governor’s Executive Budget.” 33 out of 46 respondent states claim that performance affects funding in the next fiscal year, and 11 of the 46 claim that performance affects funding even in the current fiscal year (Jordan & Hackbart, 1999, p. 75, 77). Note: survey did not give respondents opportunity to indicate degree of impact on budget allocations.
- Poister and Streib, Date: 1997; Level: Local.
  - Nature of study: Survey of city government. 695 respondents, most city managers or assistant managers.
  - Findings: 60 percent of respondents from jurisdictions with “centralized, citywide performance measurement systems that incorporate most departments and programs” reported moderate or substantial changes in budget allocations as an impact of performance measures (Poister and Streib, 1999, pp. 332–33).
- GAO, Date: 2000; Budget context: Internal Agency Budgeting (I); Level: National.
  - Nature of study: Survey of sample of managers and supervisors in 28 agencies of the federal government. 2,510 useable responses.
  - Findings: 43 percent of respondents reported using performance measures to a “great” or “very great” extent in allocating resources. A similar percentage (45 percent) reported making use of performance measures for priority setting (GAO, 2001b, pp. 11, 29-32).

*Italic: Source content from _wp05210 (Section IV and Appendix I) as provided.*

### 26.3 percent of respondent state budget offices affirmed that output or outcome

### _wp05210 - 26.3 percent of respondent state budget offices affirmed that output or outcome

### Survey findings on use of performance measures in resource allocation decisions
- 26.3 percent of respondent state budget offices affirmed that output or outcome measures were either “very important” or “important” in decisions about budget appropriations to agencies.
- 35.1 percent of respondent state budget offices believed that output or outcome measures were important or very important in agency “budget development.”
- 28.1 percent of respondent state budget offices believed that these measures were important or very important in agency “budget execution.”
- Corresponding city/county figures cited:
  - 38.7 percent of city/county respondents affirmed output or outcome measures were “very important” or “important” in decisions about budget appropriations to agencies.
  - 46.0 percent believed output or outcome measures were important or very important in agency “budget development.”
  - 35.9 percent believed these measures were important or very important in agency “budget execution.”
- Reported perception of inadequate weight of performance measurements:
  - 40 percent of state budget offices, 25.4 percent of state agency officials, and 33.9 percent of city/county respondents viewed “performance measurements” not carrying enough weight in budgetary decisions’ as a “significant problem.” (Melkers, Willoughby, James, Fountain and Campbell, 2002, p. 29).
- Other survey results cited:
  - GAO (1997b): 21 percent of respondents reported that performance measures were used to a “great” or “very great” extent in developing agency budgets; 20 percent reported similar use as the basis for funding decisions.
  - Jordan & Hackbart (1997): 23 of 41 respondent states “agree” with the proposition that “performance indicators are an important tool for making budget allocation decisions in my state.” 3 of the 41 “strongly agree”.
  - Poister and Streib (1997): Almost two-thirds of respondents from jurisdictions with “centralized, citywide performance measurement systems that incorporate most departments and programs” believed performance measures were important or very important for budgeting purposes.
  - GASB directed survey (Melkers and Willoughby, 1996): Of state government departments responding, 44.7 percent claimed to use output indicators and 41.7 percent claimed to use outcome indicators for resource allocation purposes. For municipalities the figures were 20.1 percent (output) and 17.7 percent (outcome). For counties the figures were 33.6 percent (output) and 28.1 percent (outcome). For the entire sample of 900 public sector entities the figures were 27.7 percent (output) and 25.2 percent (outcome).
    - Recalculated as a percentage of those claiming to have developed outcome/output measures:
      - state government departments—68.6/64.0 percent;
      - municipalities—66.1/58.1 percent;
      - counties—84.5/70.7 percent.
  - Wang (1996): Survey of Florida municipal officials (178 responses) — average rating of the role of performance measures in “making budget allocation decisions in setting up funding priority and funding levels” was 3.7 (3.7 from finance directors) on a 1–5 scale.

### Time-series and trends in state use of outcome/output information
- Lee (1970–1995, at 5-yearly intervals; state-level executive budget offices):
  - The percentage of states in which the state budget offices claimed that “substantial use” was made of outcome/output information in the formulation of the executive budget:
    - 15/15 percent in 1970
    - 26/45 percent in 1990
    - 18/30 percent in 1995
  - The percentage of states in which the state budget offices believed that “substantial use” was made of outcome/output information in legislative budget decision-making:
    - 25/19 percent in 1970
    - 11/10 percent in 1995
  - The percentage of states claiming they made “substantial use” or “some degree” of program analysis in the formulation of the executive budget rose until 1990, and then declined between 1990 and 1995:
    - effectiveness analysis dropped from 90 percent to 72 percent;
    - productivity analysis dropped from 100 percent to 84 percent.
  - Note: it is not possible to disaggregate the “substantial use” respondents from the “some degree” of use respondents; numbers of “substantial use” respondents are significantly smaller than “some degree.”

### GASB case-study findings on use and effects of performance measures for budgeting, management, and reporting (selected state and local summaries)
- Overall patterns:
  - Several states and localities had long-standing performance measurement efforts but mixed evidence on direct links to resource allocation.
  - Where performance budgeting or managing-for-results systems were legislated or institutionalized, reported use and effects varied from “little” or “unclear” to “substantial.”
- Selected state summaries:
  - Maine (Tucker and Campbell, 2002): Performance budgeting legislated 1997 for implementation by 2001; “had not occurred at the time of this study.” Little demonstrable effect on resource-allocation process at time of study.
  - Wisconsin (Melkers and Mhatre, 2002): Performance measurement used for “many years”; efforts to use them in budget process recent (pilot 1999); interviewees saw “little evidence of an impact on resource allocation.”
  - Arizona (Tucker, 2000a): Legislation requiring performance measurement passed in 1993; attempt to integrate with budgeting in 1999-2001 biennium; “not one individual provided an example of a direct relationship between the use of performance measures and funding levels.”
  - Illinois (Tucker, 2000b): Initiative relatively new; little use; “There has not been a measurable result in terms of dramatic allocation changes or trends realized.”
  - Iowa (Epstein and Campbell, 2000a): “Budgeting for Results” introduced in 1995; agencies used budgeting for results and performance measurement to develop and justify budgets and to target allocation of resources after they are budgeted; governor uses information in making budget decisions; “In some cases, it has resulted in a reallocation of resources.”
  - Louisiana (Epstein and Campbell, 2000b): 1997 legislation required a “managing for results cycle”; “inclusion of performance information in budget proposals...had affected the budget process”; influenced decisions on proposed new programs or funding increases and decreases.
  - Oregon (Fountain, 2000): Longstanding indicator development; performance measures provide elected officials with information to assist decision making rather than directly setting funding levels; significant contribution to performance improvement.
  - Texas (Tucker, 2000): Considerable effort since 1991; “Used extensively.” Case study indicates “Substantial” use and effects on performance improvement.
- Selected city and county summaries:
  - Austin (Epstein and Campbell, 2000c): New managing-for-results system being implemented; “Performance information is used on several levels in Austin for resource allocation.” Effects more likely driven by departmental uses than Council uses; cited explicit cost savings and service performance improvements.
  - Multnomah County, Oregon (Bernstein, 2000a): Performance budgeting since 1993–94; “not clear whether decisions ...are really driven by the use of performance measures”; nonetheless “overall ...positive effect on the performance of County programs and the achievement of goals and outcomes.”
  - Portland, Oregon (Bernstein, 2000b): Substantial measurement efforts over many years; use in budgeting apparently only “marginally”; substantial use for performance improvement.
  - Sunnyvale, California (Epstein, Campbell and Tucker, 2002a): “consistent, systematic use of performance measures for over twenty years” for resource management and continual performance improvement; Substantial use and concrete examples of meeting officials’ intentions for resource allocation; claimed to have boosted productivity and service quality substantially.
  - Winston-Salem (Bernstein, 2000d): Performance measurement used for 25 years; “performance measurement ...helped managers better allocate resources”; saved money and allowed government to do more with less.
  - Prince William County, Virginia (Bernstein, 2002): Incremental managing-for-results conversion throughout the 1990s; majority interviewed indicated performance measures are considered in terms of prior-year results and targets; cited decline in overall cost of government and shifting of resources to strategic initiatives between 1992 and 1998 as evidence that outcome budget process is working.

### Key takeaways
- Reported importance of output/outcome measures in budget appropriations varies across surveys and government levels, with figures in the source ranging from 21 percent to 46.0 percent depending on question, respondent type, and jurisdiction.
- Multiple studies report that performance measurement systems can assist budget formulation, justify budgets, and influence allocation decisions, but direct, consistent links between performance measures and funding levels are often limited or unclear in practice.
- Jurisdictions with institutionalized or long-standing performance systems (examples include Texas, Sunnyvale, Portland, Austin, Iowa) show more evidence of use and some reported effects on resource allocation and performance improvement, though effects are heterogeneous across cases.

*Source: _wp05210 - 26.3 percent of respondent state budget offices affirmed that output or outcome*

### References

### _wp05210 - References

### Performance budgeting, management reforms, and government reports
- Diamond, Jack, 2003a, Performance Budgeting: Managing the Reform Process, IMF Working Paper WP/03/33.
- Diamond, Jack, 2003b, From Program Budgeting to Performance Budgeting: The Challenge for Emerging Market Economies, IMF Working Paper WP/03/169.
- Office of Management and Budget, 2003, Assessing Program Performance for the FY 2004 Budget. Available at http://www.whitehouse.gov/omb/budintegration/part_assessing2004.html. Accessed May 14, 2003.
- Executive Office of the President/Office of Management and Budget, 2002, The President’s Management Agenda, Fiscal Year 2002, Washington: OMB.
- General Accounting Office, 1993a, Performance Budgeting: State Experiences and Implications for the Federal Government, GAO/AFMD–93–41.
- General Accounting Office, 1994a, Managing For Results: State Experiences Provide Insights for Federal Management Reforms, GAO/GGD–95–22.
- General Accounting Office, 1995a, Managing for Results: Experiences Abroad Suggest Insights for Federal Management Reforms, GAO/GGD–95–120.
- General Accounting Office, 1997a, Performance Budgeting: Past Initiative Offer Insights for GPRA Implementation, GAO/AOMD-97-46.
- General Accounting Office, 1997b, The Government Performance and Results Act: 1997 Governmentwide Implementation will be Uneven, GAO/GGD–97–109.
- General Accounting Office, 1999a, Performance Budgeting: Initial Experiences Under the Results Act in Linking Plans with Budgets, GAO/AIMD/GGD–99–67.
- General Accounting Office, 2001a, Results-Oriented Budget Practices in Federal Agencies, GAO–1084SP.
- General Accounting Office, 2001b, Managing for Results: Federal Managers’ Views on Key Management Issues Vary Widely Across Agencies, GAO–01–592.
- General Accounting Office, 2004, Observations on the Use of OMB’s Program Assessment Rating Tool for the Fiscal Year 2004 Budget, GAO–04–174.
- Government Performance Project, 2003, Paths to Performance in State & Local Government: a Final Assessment from the Maxwell School of Citizenship and Public Affairs. Accessed 13 January 2004 at ttp://www.maxwell.syr.edu/gpp/about/facts.asp.
- National Audit Office, 2001, Measuring the Performance of Government Departments, London: The Stationery Office.
- ANAO (Australian National Audit Office), 1997, Program Evaluation in the Australian Public Service, Audit report No. 3, 1997–98) (Canberra: Australian Government Publishing Service).

### Case studies, GASB research, and subnational experiences
- Bernstein, David J., 2000a, GASB SEA Research Case Study: Multnomah County, Oregon: A Strategic Focus on Outcomes (city/: Government Accounting Standards Board.
- Bernstein, David J., 2000b, GASB SEA Research Case Study: Portland, Oregon: Pioneering External Accountability, Government Accounting Standards Board.
- Bernstein, David J., 2000c, GASB SEA Research Case Study: Tucson, Arizona: An Evolving Performance Measurement Culture, Government Accounting Standards Board.
- Bernstein, David J., 2000d, GASB SEA Research Case Study: City of Winston-Salem, North Carolina: Focusing on Government Efficiency and Public Confidence, Government Accounting Standards Board.
- Bernstein, David J., 2002, GASB SEA Research Case Study: Prince William County, Virginia—Developing a Comprehensive Managing-for-Results Approach, Government Accounting Standards Board.
- Epstein, Paul D. and Wilson Campbell, 2000a, GASB SEA Research Case Study: Iowa, Government Accounting Standards Board.
- Epstein, Paul D. and Wilson Campbell, 2000b, GASB SEA Research Case Study: State of Louisiana, Government Accounting Standards Board.
- Epstein, Paul D. and Wilson Campbell, 2000c, GASB SEA Research Case Study: City of Austin, Government Accounting Standards Board.
- Epstein, Paul D., Wilson Campbell, and Laura Tucker, 2002a, GASB SEA Research Case Study: City of Sunnyvale, California, Government Accounting Standards Board.
- Epstein, Paul D., Wilson Campbell, and Laura Tucker, 2002b, GASB SEA Research Case Study: City of San Jose, Government Accounting Standards Board.
- Fountain, Jay, 2000, GASB SEA Research Case Study: State of Oregon: a Performance System Based on Benchmarks, Government Accounting Standards.
- Melkers, Julia and Pratrik Mhatre, 2002, Case Study: Wisconsin. Use and Effects of Using Performance Measures for Budgeting, Management and Reporting, Government Accounting Standards Board.
- Melkers, Julia E., Katherine G. Willoughby, Brian James, Jay Fountain, Wilson Campbell, 2002, Performance Measurement at the State and Local Levels: A Summary of Survey Results, GASB.
- Tucker, Laura, 2000a, GASB SEA Research Case Study: State Of Arizona—Focus On Performance, Government Accounting Standards Board.
- Tucker, Laura, 2000b, GASB SEA Research Case Study: State Of Illinois—Emphasis on Accountability and Managing for Results, Government Accounting Standards Board.
- Tucker, Laura, 2000c, GASB SEA Research Case Study: State Of Texas, Government Accounting Standards Board.
- Tucker, Laura and Katherine Willoughby, 2002, GASB SEA Research Case Study: DeKalb County, Georgia, Government Accounting Standards Board.
- Tucker, Laura and Wilson Campbell, 2002, Case Study: Maine. Use and Effects of Using Performance Measures for Budgeting, Management and Reporting, Government Accounting Standards Board.
- GASB (Government Accounting Standards Board), 2000, Performance Measurement for Government: State and Local Government Case Studies on Use and the Effects of Using Performance Measures for Budgeting, Management, and Reporting, Available at: http://www.seagov.org/sea_gasb_project/case_studies.shtml

### Health care financing, prospective payment systems, and casemix studies
- Averill, Richard F., Michael J. Kalison, James C. Vertrees, and Norbert I. Goldfield, 1996, “Achieving Short-Term Medicare Savings Through the Expansion of the Prospective Payment System,” Health Care Management Review, Vol. 21, No. 4, pp.18–25.
- Bray, Nancy, Carol Carter, Allen Dobson, Michael J Watt and Stephen Shortell, 1996, “An Examination of Winners and Losers under Medicare’s Prospective Payment System,” Health Care Management Review, Volume 19, No. 1, p. 44 passim.
- Boccuti, Christina and Marilyn Moon, 2003, “Comparing Medicare and Private Insurers: Growth Rates in Spending Over Three Decades,” Health Affairs, Vol. 22, No. 2, pp 230–37.
- Coulam, Robert F and Gary L Gaumer, 1991, “Medicare’s Prospective Payment System: A Critical Appraisal,” Health Care Financing Review, 1991 Annual Supplement, pp. 45–77.
- Manton, Kenneth G, Max A. Woodbury, James C. Vertrees and Eric Stallard, 1993, “Use of Medicare Services Before and After Introduction of the Prospective Payment System,” Health Services Research, 28(3), pp. 269–93.
- Rosenberg, Marjorie A. and Mark J. Browne, 2001, “The Impact of the Inpatient Prospective Payment System and Diagnosis-Related Groups: A Survey of the Literature,” North American Actuarial Journal, Vol. 5, No. 4, pp. 84–94.
- Coulam, Robert F and Gary L Gaumer, 1991, “Medicare’s Prospective Payment System: A Critical Appraisal,” Health Care Financing Review, 1991 Annual Supplement, pp. 45–77.
- Coffey, Rosanna M., 1999, “Casemix Information in the United States: Fifteen Years of Management and Clinical Experience,” Casemix Quarterly, Vol. 1, No. 1.
- Dismuke, C. E. and P. Guimaraes, 2002, “Has the Caveat of Case-Mix Influenced the Quality of Inpatient Hospital Care in Portugal?,” Applied Economics, Vol. 24, pp. 1301–1307.
- Dismuke, Clara Elizabeth and Vania Sena, 1999, “Has DRG Payment Influenced the Productive efficiency and Productivity of Diagnostic Technologies in Portuguese public hospitals? An Empirical Analysis Using Parametric and Non-Parametric Methods,” Health Case Management Science, Vol. 2, pp. 107–116.
- Palmer, George and Beth Reid, 2001, “Evaluation of the Performance of Diagnosis-Related Groups and Similar Casemix Systems: Methodological Issues,” Health Services Management Research, 14, pp. 71–81.
- Worzala, Chantal, Julian Pettengill and Jack Ashby, 2003, “Challenges and Opportunities for Medicare’s Original Prospective Payment System,” Health Affairs, Vol. 22, No. 6, pp. 175–182.
- Berwick, Donald M., et. al., 2003, “Paying for Performance: Medicare should Lead,” Health Affairs, Vol. 22, No. 6, pp. 8–9.
- Salber, Patricia R. and Bruce E. Bradley, 2002, “Adding Quality to the Health Care Purchasing Equation,” Health Affairs.
- Rosenberg, Marjorie A. and Mark J. Browne, 2001, “The Impact of the Inpatient Prospective Payment System and Diagnosis-Related Groups: A Survey of the Literature,” North American Actuarial Journal, Vol. 5, No. 4, pp. 84–94.
- Wiley, Miriam M., 1992, “Hospital Financing Reform and Case-Mix Measurement: an International Review,” Health Care Financing Review, Vol. 13, No. 4, pp. 119 passim.

### Incentives, motivation, and performance measurement theory
- Bénabou, Roland and Jean Tirole, 2003, “Intrinsic and Extrinsic Motivation,” The Review of Economic Studies, Vol. 70, pp. 489–520.
- Frey, Bruno S., 1994, “How Intrinsic Motivation is Crowded In and Out,” Rationality and Society, Vol. 6, No. 3, pp. 334–52.
- Frey, Bruno S., 1997a, Not Just for the Money: an Economic Theory of Personal Motivation, Cheltenham: Edward Elgar.
- Frey, Bruno S., 1997b, “On the Relationship between Intrinsic and Extrinsic Work Motivation,” International Journal of Industrial Organization, Vol.15, No. 4, pp. 427 passim.
- Frey, Bruno S. and Felix Oberholzer-Gee, 1997, “The Cost of Price incentives: An Empirical Analysis of Motivation Crowding-Out,” The American Economic Review; Vol. 87, No. 4, pp. 746–55.
- Dixit, Avinash, 2002, “Incentives and Organizations in the Public Sector: an Interpretive Review,” Journal of Human Resources, Vol. 37, No. 4, pp. 696–727.
- Prendergast, Canice, 1999, “The provision of incentives in firms,” Journal of Economic Literature, Vol. 37, pp.7–63.
- Gibbons, Robert, 1998, “Incentives in Organizations,” The Journal of Economic Perspectives, 12(4), pp. 115–32.
- Kreps, David M., 1997, “Intrinsic Motivation and Extrinsic Incentives,” American Economic Review, Vol. 87, No. 2; pp. 359–64.
- Heckman, James, Carolyn Heinrich, and Jeffrey Smith, 1997, “Assessing the Performance of Performance Standards in Public Bureaucracies,” American Economic Review, Vol. 87, No. 2, pp. 389–95.
- Besley, Timothy and Maitreesh Ghatak, 2003, “Incentives, Choice, and Accountability in the Provision of Public Services,” Oxford Review of Economic Policy, Vol. 19, No. 2, pp. 235–49.
- Burgess, Simon and Marisa Ratto, 2003, “The Role of Incentives in the Public Sector: Issues and Incentives, Oxford Review of Economic Policy, Vol. 19, No. 2, pp. 285–300.
- Crewson, Philip E., 1997, “Public-Service Motivation: Building Empirical Evidence of Incidence and Effect,” Journal of Public Administration Research and Theory, Vol. 7, No. 4, pp. 499–518.
- Wright, Bradley, 2001, “Public Sector Work Motivation: a Review of the Current Literature and a Revised Conceptual Model,” Journal of Public Administration Research and Theory, Vol. 11, No. 4, pp. 559–86.
- Taylor, Paul J. and Jon L. Pierce, 1999, “Effects of Introducing a Performance Management System On Employee”s Subsequent Attitudes and Effort,” Public Personnel Management, Vol. 28, No. 3, pp. 423–52.
- Deckop, John, Robert Mangel and Carol Cirka, 1999, “Getting More Than You Pay For: Organizational Citizenship Behavior and Pay-For-Performance Plans,” Academy of Management Review, Vol. 42, No. 4, pp. 420–28.
- Behn, Robert D., 1994, “The Wrong Way to Motivate,” Governing, Vol. 8 (December), p. 70.

### Budgetary theory, incrementalism, and public finance
- Berry, William D., 1990, “The Confusing Case of Budgetary Incrementalism: Too Many Meanings for a Single Concept,” Journal of Politics, Vol. 52, No.1, pp. 167–96.
- Dezhbakhsh, Hasem, Soumaya M. Tohamy and Peter H. Aranson, 2003, “A New Approach for Testing Budgetary Incrementalism,” The Journal of Politics, Vol. 65, No. 2, pp. 532–558.
- Wildavsky, Aaron B., 1984, The Politics of the Budgetary Process, Boston: Little Brown.
- Rubin, Irene S, 1997, The Politics of Public Budgeting: Getting and Spending, Borrowing and Balancing, 3rd edition, Chatham House Publishers, Chatham, NJ.
- Tanzi, Vito and Ludger Schuknecht, 2000, Public Spending in the 20th Century, Cambridge: Cambridge University Press.
- Grout, Paul A. and Margaret Stevens, 2003, “The Assessment: Financing and Managing Public Services,” Oxford Review of Economic Policy, Vol.19, No. 2, pp. 215–34.
- Robinson, Marc, 2000, “Contract Budgeting,” Public Administration, Vol. 78, No. 1, pp. 75–91.
- Robinson, Marc, 2002a, “Financial Control in Australian Government Budgeting,” Public Budgeting and Finance, Vol. 22., No.1, pp. 80–94.
- Robinson, Marc, 2002b, “Output-Purchase Funding and Budgeting Systems in the Public Sector,” Public Budgeting and Finance, Vol. 22, No. 4, pp. 17–33.
- Blondal, Jon, 2003, “Budgeting in the United States,” OECD Journal on Budgeting, Vol. 3, No. 2, pp. 7–54.
- Pyhrr, Peter A., 1973, Zero-Base Budgeting: a Practical Management Tool for Evaluating Expenses, New York: John Wiley & Sons.
- Mikesell, John L, 1995, Fiscal Administration: Analysis and Applications for the Public Sector, 4th edition, Wadsworth; Belmont, CA.

*References list as presented in _wp05210 - References (source PDF).*

---


_Source: https://www.imf.org/-/media/websites/imf/imported-full-text-pdf/external/pubs/ft/wp/2005/_wp05210.pdf_
