Geology ReportsSearch

USGS · 70209617

Well predictive performance of play-wide and Subarea Random Forest models for Bakken productivity

Abstract

In recent years, geologists and petroleum engineers have struggled to clearly identify the mechanisms that drive productivity in horizontal, hydraulically-fractured oil wells producing from the middle member of the Bakken formation. This paper fills a gap in the literature by showing how this play’s heterogeneity affects factors that drive well productivity. It is important because understanding the relative strength of productivity drivers and how predictors vary spatially facilitates best-practices for well site selection and well completion design. The paper describes an application of the Random Forest (RF) machine learning technique to identify these mechanisms and to evaluate their importance across 9 subareas of the North Dakota portion of the Bakken play. The study examined productivity of 7311 wells initiating production from 2010 through 2017. Well productivity varied considerably across the 9 subareas within the play, so it was not surprising that the dominant predictors, the initial 180-day water cut and the 30-day initial gas production, vary spatially to mirror local conditions that strongly affect well productivity. The relative importance of well completion predictor variables, that is, the numbers of fractures stages per well, volume of injected proppant per stage, volume of injected fluids per stage, and lateral length, varied considerably across the subareas. Statistical permutation tests are presented that generally confirm the importance rankings. Subarea Random Forest models explained from 50 percent to 82 percent of the variation in productivity test samples while the play-wide model explained 73 percent of the test sample well productivity. Weakness in the predictive ability of the Random Forest models are traced to the limited variability in the training data. Implications of the empirical findings regarding the Bakken play for operators and for research and government institutions are discussed in the concluding section.

Explore related subjects

90° N90° S · 180° W ← longitude → 180° E
Source-reported bounding extent: 45.02695045318546° to 48.980216985374994° latitude; -109.072265625° to -98.701171875° longitude. This indicates report coverage, not an exact sampling location. View area on OpenStreetMap.

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Emil D. Attanasi, Philip A. Freeman, Tim Coburn. 2020. Well predictive performance of play-wide and Subarea Random Forest models for Bakken productivity. https://doi.org/10.1016/j.petrol.2020.107150

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related USGS reports

Oil-source rock correlation studies in the unconventional Upper Cretaceous Tuscaloosa marine shale (TMS) petroleum system, Mississippi and Louisiana, USA

The U.S. Geological Survey assessed undiscovered unconventional hydrocarbon resources reservoired in the Upper Cretaceous Tuscaloosa marine shale (TMS) of southern Mississippi and adjacent Louisiana in 2018. As part of the assessment, oil-source rock correlations were examined in the TMS play area where operators produce light (38–45° API), sweet oil from horizontal, hydraulically-fractured wells in an overpressured ‘high-resistivity’ (>5 Ω-m) zone at the base of the TMS. Geochemical data from 39 oil samples and 17 source rock solvent extracts collected from the TMS play area indicate close correspondence for Tuscaloosa Group oils [from lower Tuscaloosa, middle Tuscaloosa (the TMS) and upper Tuscaloosa reservoirs] in thermal maturity (computed from MPI), SARA proportions, n- alkane distributions, isoprenoid and DBT/P ratios, monoaromatic steroids, and δ 13 C isotopic compositions (from whole oils, saturate and aromatic fractions). Other parameters (normal steranes, extended homohopanes, C 31 R/C 30 hopane, norhopane/hopane and tricyclic terpane ratios, gammacerane/hopane) show most oil samples have similar values, suggesting all Tuscaloosa Group oils are from a common mixed marine-terrigenous source rock. Tighter distributions for triaromatic steroid (TAS) and δ 13 C isotopic composition for conventional oils in lower and upper Tuscaloosa reservoirs may indicate charge occurred in a single or shorter pulse relative to TMS oils which show broader TAS and δ 13 C properties, possibly from their generation over an extended period of burial maturation. Dissimilarity in geochemical properties between lower Tuscaloosa source rock solvent extracts and Tuscaloosa Group oils indicates lower Tuscaloosa source rocks did not contribute significantly to conventional and unconventional Tuscaloosa Group hydrocarbon accumulations. Whereas, TMS solvent extracts are similar to Tuscaloosa Group oils, suggesting an oil-source rock correlation. Excluding the possibility for long-distance lateral migration from a similar source downdip (which is unnecessary given thermal maturity considerations), the observations indicate 1. the TMS is a self-sourced reservoir, 2. the TMS is the source of oils accumulated in nearby conventional Tuscaloosa Group reservoirs, and 3. thin organic-rich shales in the lower Tuscaloosa did not contribute substantially to any oil accumulations in the Tuscaloosa Group.

Louisiana, Mississippi

An ANCOVA model for porosity and its uncertainty for oil reservoirs based on TORIS dataset

Porosity is one of the most important parameters to assess in-place oil or gas in reservoirs, and to evaluate recovery from enhanced production operations. Since it is relatively well-established to determine porosity using different laboratory and field methods, its value is usually determined at many locations across a reservoir as part of the common practice to capture reservoir heterogeneity and the variability in values. This suite of measurements and the distribution of values are most valuable for probabilistic reservoir assessments, and for spatial modeling if the exact data locations are known. Despite the importance of individual measurements to set the range of values for probabilistic studies, it is not always possible to access these data due to confidentiality. In most cases, commercial or publicly available databases that assessments may rely on usually report only mean values of porosity, like any other reservoir data, or they may not report a value at all. This makes both quantifying the mean value and the uncertainty around it difficult for probabilistic assessments. In this study, the TORIS (Tertiary Oil Recovery Information System) dataset of the National Petroleum Council and the U.S. Department of Energy was used to model porosity and the uncertainty around predicted values. TORIS is an integrated dataset of production data, reservoir properties, and project databases of crude oil reservoirs in the United States. The model presented in the paper was based on ANCOVA (Analysis of Co-Variance) of data from 1038 reservoirs from the TORIS dataset for porosity prediction, validation and testing for quantitative and qualitative parameters that may be readily available in most cases, and to estimate uncertainty around the mean values. This model also explored association of porosity values to different parameters, and to different depositional systems and diagenetic overprint conditions. Furthermore, an ANN (Artificial Neural Network) model was created to compare the predicted values of both models. Results showed that the ANN model was able to represent more of the variability, however it lacked the insights that might be gained from the ANCOVA model.

Journal of Petroleum Science and Engineering

A fuzzy logic approach for estimating recovery factors of miscible CO2-EOR projects in the United States

"Recovery factor (RF) is one of the most fundamental parameters that define engineering and economical success of any operational phase in oil and gas production. The effectiveness of the operation, e.g. CO2-EOR (enhanced oil recovery with carbon dioxide injection), is usually defined by multiplying the resultant recovery factor by the original oil in place. Moreover, investment decisions for such engineering projects are also performed based on predicted recovery factors. Despite its importance, though, it is not easy to predict recovery factors as they are affected by many factors including the type of the recovery process, reservoir type, fluid properties, reservoir heterogeneity, depth, thickness, to name a few. The usual method of estimating recovery factors is laboratory experiments or numerical modeling, each of which has their own limitations due to data requirements, boundary conditions and scale effects. In this work, a fuzzy inference system approach has been adopted to predict miscible CO2-EOR recovery factors of the major field applications in the United States with the premise that it can be used as a guidance tool for making decisions based on different inputs. The fuzzy system was build using a Mamdani-type fuzzy logic inference engine, and by using reservoir data compiled from different sources as inputs and recovery factors gathered from a literature survey. Due to the limited number of field cases that could be used for this purpose, 24 sets of applications were included in the study. Selected input variables were water saturation after waterflood (Sorw), well spacing, porosity, permeability, depth, net pay thickness, initial pressure, API gravity of oil, hydrocarbon pore volume CO2 injected, and reservoir lithology. The type of membership functions were decided based on the system’s predictive performance. The model showed reasonable predictive capability for the field observations of recovery factor despite the complexity of this parameter. In addition, since the fuzzy solution was multi-dimensional due to multiple inputs, system behavior was used to demonstrate response of miscible CO2-EOR recovery factor to different inputs. "

Journal of Petroleum Science and Engineering