## wpiea2020057-print-pdf

## Source details

**Canonical URL:** [wpiea2020057-print-pdf](https://www.imf.org/-/media/files/publications/wp/2020/english/wpiea2020057-print-pdf.pdf)

## Other formats

- [Markdown version](/-/media/files/publications/wp/2020/english/wpiea2020057-print-pdf.pdf.md)
- [Structured JSON version](/-/media/files/publications/wp/2020/english/wpiea2020057-print-pdf.pdf.json)

---

### I. INTRODUCTION — Key findings and purpose
- World trade: over 80 percent of global merchandise trade by volume and more than 70 percent of its value is carried by the international shipping industry (UNCTAD, 2017).
- Objective: build real-time indicators of world seaborne trade using AIS (Automatic Identification System) messages emitted by vessels to provide a more immediate picture of global trade flows than existing proxies with one- to three-month lags.
- Methodology summary:
  - Transform raw AIS data into estimates of trade volume at the world, bilateral and within-country levels using machine-learning techniques.
  - Construct the Global Trade Intelligence (GTI) index from AIS-derived estimates.
- Predictive power and timeliness:
  - For some countries, GTI has the highest correlation with official data when GTI leads by one month, implying nowcasting and sometimes short-term forecasting capability.
  - Example: in countries with high data quality and a typical publication lag of around two months, GTI can calculate 3m/3m growth from AIS data that may predict quarter-on-quarter actual imports for the upcoming quarter.
  - CPB publication timing: on or around the 23rd day of month t, CPB publishes first estimates for month t-2 with a 7 to 11-week lag; effective lag often 11 to 15 weeks due to revisions.
  - Current implementation of methods in the paper produce import estimates with a 5-day lag, and export estimates with a 10-day lag.
- Performance benchmarks:
  - Japan: correlation between monthly GTI and official (CPB) growth rates up to 0.49 when official data measured with a one-month lead.
  - Euro Area: correlation up to 0.47 with a one-month lead.
  - World level: monthly pairwise correlation with official statistics nearly 0.9 in levels, and around 0.4 in quarter-on-quarter growth rates.
  - Crude oil case (sectoral benchmark): correlation in levels for global crude oil imports is 0.73 in the raw data, and 0.81 on a 3-month moving average basis; in growth rates, correlation as high as 0.47.
- Use cases illustrated: event studies including Hurricane Maria in Puerto Rico and the novel coronavirus outbreak in China.
- Novelty and contribution:
  - First paper to construct trade volume indicators relying solely on AIS data and publicly available sources.
  - Develops a transparent algorithm to identify port calls benchmarked against four years of daily official vessel entrances recorded by U.S. customs.
  - First to report accuracy of AIS-derived port call data with respect to official vessel-level statistics.
- Limitations and caveats:
  - Method does not work as well for all countries where shares of trade by sea are low or port geography/infrastructure differs substantially.
  - Draught information is manually entered by crews and may have a lag; consequently export estimates are censored by 10 days and imports by 5 days in current implementation.
  - Work presented as research in progress with listed ongoing methodological refinements and future research avenues.

### II. AIS DATA — Data coverage and properties
- AIS fields:
  - Manually entered: draught, destination, navigational status (e.g. moored, anchored, under way using engine).
  - Automatically generated: position (latitude and longitude), speed.
- Dataset used:
  - Provider: MarineTraffic.
  - Coverage period: January 1, 2015 to April 18, 2020.
  - Size: over one billion messages from over 50,000 distinct ships.
  - Purchased data frequency: down-sampled to hourly frequency (original messages often every 2-10 seconds).
  - Average coverage after technical limitations: about one message per ship every two hours on average.

### III. ELICITING PORTS FROM RAW AIS DATA — Unsupervised learning approach
- Problem framing:
  - Identify AIS messages occurring within port boundaries where port boundaries vary and change over time.
  - Focus on AIS messages with speed < 0.5 knots and navigational status = moored; for certain large vessel types include status = anchored with speed < 0.5 knots.
- Vessel-size exceptions (anchored messages included for these):
  - Bulk carriers with deadweight tonnage (DWT) over 75,000 metric tons.
  - Oil/chemical tankers with DWT over 50,000 metric tons.
  - Crude oil tankers with DWT over 100,000 metric tons.
- Clustering algorithm:
  - DBSCAN applied to set D of candidate messages; clusters converted to polygons via convex hull.
  - DBSCAN parameters and definitions preserved exactly:
    - ε-neighborhood: Nε(p) = { q ∈ D | geodesic distance(p,q) ≤ ε }.
    - Core point: at least N points in ε-neighborhood.
    - Directly reachable, reachable, cluster, and outlier per DBSCAN definitions.
- Rationale for DBSCAN:
  - Captures arbitrary-shaped clusters, does not require predetermined number of clusters, and identifies outliers — advantages over K-means and hotspot/spatial Poisson approaches.
- Limitations:
  - Requires exogenous choices of ε and N; parameters set to err on side of producing too many polygons to avoid missing low-traffic ports.
  - Recent extensions to endogenize ε and N proposed for future refinements.
- Implementation note: proceed to implement weighted DBSCAN and combine ML-identified potential port-call polygons with U.S. vessel entrance records in a supervised learning step.

### Data preprocessing and spatial aggregation
- Initial filtering criteria:
  - Messages with navigational status "anchored" and speed below 0.5 knots per hour are used to identify anchored port calls.
- Data reduction:
  - From an original 4-year dataset of 854.7 million observations, filtering leaves 189.0 million observations.
- Latitude/longitude rounding:
  - For latitudes between 0 and 45 degrees, round to the nearest fourth decimal place (around 11 meters at the equator and around 8 meters at 45-degree latitude).
  - For latitudes greater than 45 degrees, round to the nearest one-fifth of the fourth decimal place (i.e. in 0.0002 increments).
- Position weighting:
  - Dataset collapsed by rounded-up latitude and longitude, creating position_weight that counts the number of AIS messages at each rounded-up position.
- Weighted DBSCAN parameters:
  - Weighted DBSCAN used with each grid point weighted by number of messages.
  - Parameters set: ε = 2,000 meters and N = 1,000.
  - Calibration rationale for N: a single-berth terminal must be occupied at least about five percent of the time to be identified; with 12 daily messages over four years a constantly occupied berth would have 12x365x4 = 17,520 messages; DBSCAN parameter set to 1,000.
- Sensitivity and distances:
  - Sensitivity: number of noise points does not change dramatically under different ε choices.
  - Distances computed using the haversine formula.

### Assigning AIS-derived polygons to countries and ports
- Centroid-to-country assignment:
  - Use version 10 of the Maritime Boundaries Geodatabase (Flanders Marine Institute, 2018) to map centroids to sovereign countries via territorial waters polygons.
  - If centroid search over water returns no sovereign country, use polygons over land areas from the World Borders Dataset.
  - Outcomes for ~3,500 cluster centroids:
    - All but around 1,300 are mapped to a country based on maritime boundaries.
    - Of those 1,300, all but around 500 are mapped to a country based on land boundaries.
    - Final result: about 3,000 centroids mapped to a country and about 500 centroids not assigned to any country.
- Assigning centroids to head ports (hierarchical two-step):
  - Step 1: If centroid has country assigned, search for head ports within a 30 km radius in same country; if non-empty assign to closest head port in set.
  - Step 2: If no head port in same country within 30 km, assign to WPI port in same country closest to centroid.
  - If centroid lacks country assignment: search for head ports within 30 km and assign to closest; otherwise assign to closest WPI port.
  - Rationale: 30-km calibration consistent with port-group construction and avoids misassigning to non-head ports.
- Removing distant assignments and synthetic ports:
  - If centroid-to-nearest-port distance > 75 km, assignment removed; such clusters termed "synthetic ports".
  - Synthetic ports that lie on waters or land of a country are synthetic ports of that country.
  - Total synthetic ports: 16, assigned to closest countries and manually validated.

### Identifying port calls – supervised learning overview
- Problem: port-polygon presence does not imply a true port call; ships may traverse port polygons without stopping to load/unload (examples: Rotterdam river channels, ships crossing San Francisco Bay to reach Oakland, Singapore Strait transits).
- Reference data for supervised labeling:
  - Official vessel entrances and clearances at U.S. ports from the U.S. Army Corps of Engineers’ Navigation Data Center (NDC), compiled with U.S. Department of Homeland Security’s Customs and Border Protection.
  - NDC records most port calls on U.S. ports and include vessels’ IMOs and date of entry; available through end 2018 as of March 2020.
- Treatment of U.S.-flagged vessels:
  - By regulation, all vessels except U.S.-flagged ships coming directly from another U.S. port and without foreign goods onboard must file an entrance statement.
  - U.S.-flagged ships make up only 0.5 percent of the dataset.
  - Analyses restrict attention to non-U.S.-flagged vessels that in raw database arrive from non-U.S. ports.
- Matching rules between AIS-derived polygon visits and NDC:
  - Matched port calls:
    - For each NDC port call, search for same IMO in port visit data over a window of +/- 2 days around the NDC-recorded date.
    - Rationale for +/- 2 days:
      - By law vessels have up to two days to file their entrance report.
      - UTC vs local U.S. time discrepancies: AIS timestamps UTC, NDC records local U.S. date; shift AIS timestamps by six hours to limit discrepancy.
      - Vessels may report entry ahead of entering the port (e.g., while waiting at anchorage).
    - When multiple port calls within the window exist, match NDC records to ones with closest dates; if tie persists, prefer the earliest record.
  - False negatives: IMOs recording port calls in NDC for which no port call is found in AIS-derived data over +/- 2-day window.
  - False positives: IMOs recording port calls in AIS-derived data that remain unmatched to NDC port calls over +/- 2-day window; note some vessels are not required to file an entrance record, so false positives do not necessarily imply non-entry.

### Mapping results and assessing blind spots (2015–2018)
- Aggregated counts and matching summary (2015–2018):
  - AIS-derived GTI polygons record more than 500,000 distinct ship-date pairs with messages coming from within polygons associated with U.S. ports.
  - NDC shows nearly 290,000 entrance records.
  - Around 215,000 NDC port calls can be mapped to a GTI polygon visit within a +/- 2-day window.
- Initial classification from matching (Table 1 reproduction preserved as in source):
  - Found by GTI on day t = -2: 80; FrequencyPercentCumulative line shows "-28040.10.1".
  - Found by GTI on day t = -1: 9,458 (1.6 percent; cumulative 1.7 percent).
  - Found by GTI on day t = 0: 189,538 (31.4 percent; cumulative 33.1 percent).
  - Found by GTI on day t = +1: 104,427 (17.3 percent; cumulative 34.8 percent).
  - Found by GTI on day t = +2: 4,460 (0.7 percent; cumulative 35.5 percent).
  - In ML-AIS but not in NDC (false positives): 318,168 (52.6 percent; cumulative 88.1 percent).
  - In NDC but not in ML-AIS (false negatives): 71,674 (11.9 percent; cumulative 100.0 percent).
  - Total observations: 604,529 (100.0 percent).
- Interpretation of false negatives:
  - Many apparent false negatives correspond to IMOs that never appear in AIS-based data (mostly small ships such as tugs, barges, service vessels not participating in international trade).
  - Of the ~71,000 false negative port calls, about 58,000 correspond to IMOs not in the AIS data.
  - Excluding these IMOs yields higher matching performance (detailed exclusion numbers lie beyond this excerpt).

### Detection performance of AIS-derived port polygons
- Coverage:
  - At 94.3 percent, the GTI polygons capture the overwhelming majority of official U.S. port calls.
  - Total labeled official port calls used in some tables: Total = 227,621 (100.0).
  - In NDC but not in ML-AIS (false negatives): 12,934 (5.71 percent of the total).
- Example false negative rates by vessel type (Table 3):
  - CONTAINER SHIP: Total official port calls = 71,598; Found by GTI = 69,599; False negative rate = 2.79 percent.
  - BULK CARRIER: Total = 34,882; Found = 31,666; False negative rate = 9.22 percent.
  - OIL/CHEMICAL TANKER: Total = 32,324; Found = 30,375; False negative rate = 6.03 percent.
  - GENERAL CARGO: Total = 20,490; Found = 19,232; False negative rate = 6.14 percent.
  - VEHICLES CARRIER: Total = 18,965; Found = 18,347; False negative rate = 3.26 percent.
  - CRUDE OIL TANKER: Total = 18,182; Found = 17,352; False negative rate = 4.56 percent.
  - SELF DISCHARGING BULK CARRIER: Total = 6,608; Found = 5,422; False negative rate = 17.95 percent.
  - LPG TANKER: Total = 5,993; Found = 5,683; False negative rate = 5.17 percent.
  - OIL PRODUCTS TANKER: Total = 4,185; Found = 3,877; False negative rate = 7.36 percent.
  - RO-RO CARGO: Total = 3,348; Found = 3,248; False negative rate = 2.99 percent.
- Notes:
  - False negative rates are below 10 percent for all presented vessel types except self-discharging bulk carriers; smaller vessels within that class are more likely missed.
  - False negative rates vary by port geography; port- and vessel-type analyses performed (examples: bulk carriers).

### False positives and supervised learning approach
- Observed patterns motivating filtering:
  - For GTI port visits with matching NDC port call, about half have a previous port visit in the U.S.; this fraction rises to nearly 70 percent for GTI visits without a matching NDC port call.
  - In GTI port visits with matching NDC port call, about 3 percent of vessels are U.S.-flagged; this fraction rises to 17 percent for GTI visits without a matching NDC port call.
- Training data filters:
  - To reduce bias, observations where the previous GTI port of call is in the U.S. and the vessel is U.S.-flagged are dropped from classifier training.
- Intuition features to distinguish true port calls from false positives include:
  - Average message speed.
  - Duration within polygon.
  - Number of anchored/moored messages.
  - Fast transits (e.g., average speed ~10 knots) and durations less than one hour unlikely to represent cargo operations.

### Random forest classifier: features, training, and performance
- Motivation: non-linear and non-monotonic relations between covariates and port-call probability call for flexible, nonparametric models.
- Classifier type and loss:
  - Random forest ensemble using Gini's diversity index; deadweight-tonnage-weighted frequencies penalize misclassification proportionally to vessel gross tonnage.
- Features (16 covariates, Table 4):
  - Anchored message ratio (Fraction of messages with status anchored over anchored+moored)
  - Total no. anchored/moored messages
  - Is anchorage polygon (1 if anchorage polygon)
  - Hours from draught change
  - Dummy previous mooring call (1 if other mooring call in past 48 hours)
  - Hours from previous mooring call
  - No. draught decrease / No. draught increase
  - Share of low-speed messages (speed < 0.5 knots)
  - No. messages
  - Average message speed
  - Average speed (distance over time from first message out to first message in polygon)
  - Duration of stay (time from first message out to first message in polygon)
  - Summer deadweight tonnage
  - Gross tonnage
  - Ship type
- Draught considerations:
  - Draught manually entered and prone to measurement error and lags.
  - Use last observed draught before entering polygon, presuming more careful reporting before port entry.
- Training and tuning:
  - Labeled port visits: 472,085.
  - Test set: randomly pick 10 percent; remaining 90 percent used for training and hyperparameter tuning with five-fold cross validation.
  - Hyperparameter tuning via Bayesian optimization.
  - Tuned hyperparameters: minimum number of points at each leaf node = 1; number of trees = 232.
  - Bootstrap fraction: 0.6 of training samples for each tree.
- Model comparison on testing sample (Table 5):
  - Accuracy: RANDOM FORESTS 91.3%, SVM 84.5%, LOGISTIC REGRESSION 84.7%.
  - Precision: RANDOM FORESTS 89.4%, SVM 80.9%, LOGISTIC REGRESSION 81.2%.
  - Recall: RANDOM FORESTS 91.1%, SVM 85.0%, LOGISTIC REGRESSION 84.8%.
- Interpretation:
  - Random forest outperforms SVM and logistic regression on accuracy, precision, and recall.
  - Precision (~89.4%) indicates relatively low false positive rate; recall (91.1%) indicates relatively low false negative rate.
  - Feature importance: top two features are anchored message ratio and duration of stay within polygon; other important predictors include number of messages from polygon, ship type, and time to next mooring polygon.

### From port calls to voyages and volume of trade
- Post-classification processing:
  - Apply trained random forest to worldwide AIS-derived polygon-visit dataset to produce estimated global port call database.
  - Construct voyages between estimated port calls; "see through" and discard certain ports to avoid misclassification of bilateral origin/destination:
    - Jurisdictions discarded: Panama and Gibraltar; two Suez Canal ports (As Suways, Bur Said (Port Said)).
- Metric tons estimation per inbound ship visit:
  - Notation and estimation formula preserved:
    - m_{dmit}^{IN} = inbound metric tons of cargo transported by ship d at time t into a port.
    - dwt_i = deadweight tonnage of ship i.
    - d_{i,mdes} = maximum (design) draught of ship i.
    - d_{i,ballast} = draught when not transporting any goods (ballast draught).
    - d_{iit} = observed draught.
    - m_{dmit}^{IN} = dwt_i × (d_{iit} − d_{i,ballast}) / (d_{i,mdes} − d_{i,ballast})
- Ballast draught imputation:
  - For each vessel, compute ratio: first percentile of observed draught over four years divided by design draught.
  - By ship type and deadweight tonnage tertile, take median of this ratio, denote mmr_{rbd}.
  - Impute vessel ballast draught as: d_{i(r,b),ballast} = mmr_{rbd} × d_{i,mdes}.
- Mapping metric tons to imports/exports/internal trade (procedure):
  1. Define outgoing cargo weight for ship visit as m_{dmit}^{OUT} = m_{dmi,t+1}^{IN}. Net cargo offloaded at port: m_{dmit}^{OFF} = m_{dmit}^{IN} − m_{dmit}^{OUT} (can be negative or positive).
  2. If ship enters a port in a new country (previous port in different country) and m_{dmit}^{OFF} > 0 → label as import; if m_{dmit}^{OFF} < 0 → do not count imports on this trip (may be internal trade).
  3. If first port of call in country had an import event, count subsequent drops in metric tons as imports until the first instance the ship starts gaining metric tons (i.e., until m_{dmi,t+(J+1)}^{OUT} < 0). Exports defined analogously by scanning backwards from last port call in the country and labeling contiguous gains as exports.
- Aggregation:
  - Group port-call events into importing, exporting, or internal trade and aggregate metric tons by country grouping, vessel type, and time period (e.g., monthly country-level import/export volume indicators, weekly bilateral crude oil trade).

### Results and benchmarking
- Macroeconomic nowcasting performance:
  - Global seasonally adjusted monthly import volume gauge correlation with official statistics (CPB): 0.88.
  - In 3m/3m growth rate terms, global correlation = 0.4.
  - World exports show similar performance; growth rate of exports somewhat weaker.
  - Median correlation across CPB sample (3-month moving averages in levels): imports = 0.68, exports = 0.46.
  - Median correlations in growth rates: imports = 0.27, exports = 0.18.
  - Benchmarking limited to 43 economies (including Euro Area) plus the world due to scarcity of monthly trade volume data.
- Sectoral analysis — crude oil (JODI comparison):
  - Global crude oil imports:
    - Correlation in raw levels: 0.73.
    - Correlation on 3-month moving average basis: 0.81.
    - Correlation in 3m/3m growth rates: 0.47.
  - Exports: tracking poorer at global level; JODI import/export discrepancies (JODI global imports exceed exports by over 15 percent) and reporting differences noted.
  - Country-level mismatches can arise where port of arrival differs from final destination (examples: Germany, Netherlands).
- Event studies and daily-frequency applications:
  - Hurricane Maria (Puerto Rico, landfall September 20, 2017): offloading vessel traffic declined by as much as 75 percent and took around 10 days to recover.
  - COVID-19 pandemic examples:
    - Port-level AIS message frequency for Port of Ningbo shows dramatic reductions during strict containment and partial rebounds by April 2020.
    - Daily estimated metric tons of exports for China (normalized to 2017–2019 avg, Lunar New Year-adjusted) show dramatic fall with resumption in early–mid March but incomplete recovery by mid-April.
  - Real-time world trade estimates (30-day moving averages, relative to 2017–2019 average) show sectoral heterogeneity: oil-related exports strong at times (consistent with storage at sea) while container and vehicle exports/imports show weak readings during production halts and demand drops.

### Ongoing developments and avenues for improvement
- Algorithmic refinements not requiring new data:
  - Upgrade port-polygon construction: test OPTICS (Ankerst et al., 1999) to find contours where ship density drops and better separate low-traffic ports from nearby high-traffic ports.
  - Predict ship destination while at sea using unstructured AIS textual info and voyage history to increase forecasting horizon from ~1 month to 2–3 months.
- Improvements requiring additional data:
  - More detailed vessel information (ship registers) to refine cargo weight estimation:
    - Current method assumes linear change in cargo weight with draught (equivalent to rectangular cuboid hull); more accurate hull form coefficients (e.g., block coefficient) would improve metric-ton calculations.
    - Direct ballast draught values from registers would be preferable to imputed values.
- Domain-specific modifications:
  - Ports with unusual geographies (oil platforms, FPSOs, riverine ports) and special operations (lightering) can cause double counting or missed events; expert knowledge and port-specific rules needed.
- Trade-weighted indices:
  - Alternative mapping coarse vessel classification to HS 4-digit codes and applying COMTRADE-based weights tested (Appendix B, Table A1).
  - Results: some economies improve, others do not; imports in levels show more unambiguous improvement than growth rates.

### Concluding remarks
- The methodology provides an end-to-end proof of concept to construct trade volume indicators relying only on AIS data and off-the-shelf machine learning.
- Achieves good fit with official trade statistics for many countries and the world in aggregate; useful for sectoral crude oil analysis and granular event studies.
- Further improvements expected from algorithmic refinements, additional vessel data, and country-specific domain knowledge.
- Expanding availability of monthly official trade volume statistics for more countries would aid further benchmarking and application to emerging and developing economies that rely heavily on seaborne trade.

*Source: wpiea2020057-print-pdf (canonical PDF content provided)*

### References .............................................................................................................

### References

### I. INTRODUCTION — Key findings and purpose
- World trade: over 80 percent of global merchandise trade by volume and more than 70 percent of its value is carried by the international shipping industry (UNCTAD, 2017).
- Objective: build real-time indicators of world seaborne trade using AIS (Automatic Identification System) messages emitted by vessels to provide a more immediate picture of global trade flows than existing proxies with one- to three-month lags.
- Methodology summary:
  - Transform raw AIS data into estimates of trade volume at the world, bilateral and within-country levels using machine-learning techniques.
  - Construct the Global Trade Intelligence (GTI) index from AIS-derived estimates.
- Predictive power and timeliness:
  - For some countries, GTI has the highest correlation with official data when GTI leads by one month, implying nowcasting and sometimes short-term forecasting capability.
  - Example: in countries with high data quality and a typical publication lag of around two months, GTI can calculate 3m/3m growth from AIS data that may predict quarter-on-quarter actual imports for the upcoming quarter.
  - CPB publication timing: on or around the 23rd day of month t, CPB publishes first estimates for month t-2 with a 7 to 11-week lag; effective lag often 11 to 15 weeks due to revisions.
  - Current implementation of methods in the paper produce import estimates with a 5-day lag, and export estimates with a 10-day lag.
- Performance benchmarks:
  - Japan: correlation between monthly GTI and official (CPB) growth rates up to 0.49 when official data measured with a one-month lead.
  - Euro Area: correlation up to 0.47 with a one-month lead.
  - World level: monthly pairwise correlation with official statistics nearly 0.9 in levels, and around 0.4 in quarter-on-quarter growth rates.
  - Crude oil case (sectoral benchmark): correlation in levels for global crude oil imports is 0.73 in the raw data, and 0.81 on a 3-month moving average basis; in growth rates, correlation as high as 0.47.
- Use cases illustrated: event studies including Hurricane Maria in Puerto Rico and the novel coronavirus outbreak in China.
- Novelty and contribution:
  - First paper to construct trade volume indicators relying solely on AIS data and publicly available sources.
  - Develops a transparent algorithm to identify port calls benchmarked against four years of daily official vessel entrances recorded by U.S. customs.
  - First to report accuracy of AIS-derived port call data with respect to official vessel-level statistics.
- Limitations and caveats noted:
  - Method does not work as well for all countries where shares of trade by sea are low or port geography/infrastructure differs substantially.
  - Draught information is manually entered by crews and may have a lag; consequently export estimates are censored by 10 days and imports by 5 days in current implementation.
  - The authors consider this work research in progress and list ongoing methodological refinements and future research avenues.

### II. AIS DATA — Data coverage and properties
- AIS: Automatic Identification System; IMO regulations require AIS transponders aboard specified vessels since end-2004.
- Manually entered AIS fields: draught, destination, navigational status (e.g. moored, anchored, under way using engine).
- Automatically generated AIS fields: position (latitude and longitude), speed.
- Dataset used:
  - Provider: MarineTraffic.
  - Coverage period: January 1, 2015 to April 18, 2020.
  - Size: over one billion messages from over 50,000 distinct ships.
  - Purchased data frequency: down-sampled to hourly frequency (original messages often every 2-10 seconds).
  - Average coverage after technical limitations: about one message per ship every two hours on average.

### III. ELICITING PORTS FROM RAW AIS DATA — Unsupervised learning approach
- Problem: identify when a ship’s AIS message comes from within port boundaries; port boundaries vary and change over time.
- Motivation: National Geospatial Intelligence Agency’s World Port Index (WPI) identifies 3,669 ports, but static point-based port definitions are insufficient for heterogeneous port shapes/sizes.
- Conceptual approach:
  - Focus on AIS messages with speed < 0.5 knots and navigational status = moored; for certain large vessel types that may load/offload without mooring, also include status = anchored with speed < 0.5 knots.
  - Vessel-size exceptions (consider anchored messages for these types):
    - Bulk carriers with deadweight tonnage (DWT) over 75,000 metric tons.
    - Oil/chemical tankers with DWT over 50,000 metric tons.
    - Crude oil tankers with DWT over 100,000 metric tons.
  - Let D denote the set of such messages; identify clusters in D based on geodesic distance.
- Clustering algorithm: DBSCAN (density-based algorithm for discovering clusters in large spatial databases with noise; Ester et al., 1996).
  - DBSCAN requires two parameters: radius ε and minimum number of points N.
  - Definitions preserved exactly from source:
    - ε-neighborhood of point p: Nε(p) = { q ∈ D | geodesic distance(p,q) ≤ ε }.
    - Core point: at least N points in its ε-neighborhood.
    - Directly reachable, reachable, cluster and outlier defined per DBSCAN.
  - For each DBSCAN cluster, define polygon as the convex hull of the cluster to produce geographical boundaries.
- Rationale for DBSCAN over alternatives:
  - K-means limitations: forces every point into a cluster and requires prior knowledge of number of clusters K.
  - Hotspot/spatial Poisson approaches require prior knowledge of port boundary shapes.
  - DBSCAN advantages: elicits clusters of arbitrary shapes, does not require number of clusters, can identify outliers.
- Limitations of DBSCAN:
  - Requires exogenous choices of ε and N.
  - To avoid missing less-frequented ports, parameters are set to err on the side of producing too many polygons.
  - Recent extensions to endogenize ε and N noted as promising for future refinements.
- Implementation note (lead into next section): authors proceed to implement a weighted DBSCAN and subsequently combine ML-identified potential port-call polygons with U.S. vessel entrance records in a supervised learning step to classify actual port calls.

*Source: Excerpt from the provided PDF content unit.*

### 0.5 knots per hour. To identify port calls of certain ship types that may engage in trade while

### wpiea2020057-print-pdf - 0.5 knots per hour. To identify port calls of certain ship types that may engage in trade while

### Data preprocessing and spatial aggregation
- Initial filtering criteria:
  - Messages with navigational status "anchored" and speed below 0.5 knots per hour are used to identify anchored port calls.
- Data reduction:
  - From an original 4-year dataset of 854.7 million observations, filtering leaves 189.0 million observations.
- Latitude/longitude rounding:
  - For latitudes between 0 and 45 degrees, round to the nearest fourth decimal place (around 11 meters at the equator and around 8 meters at 45-degree latitude).
  - For latitudes greater than 45 degrees, round to the nearest one-fifth of the fourth decimal place (i.e. in 0.0002 increments).
- Position weighting:
  - The dataset is collapsed by rounded-up latitude and longitude, creating a variable position_weight that counts the number of AIS messages corresponding to each rounded-up position.

### Weighted DBSCAN clustering: parameters and rationale
- Algorithm and parameters:
  - Weighted DBSCAN is used, with each grid point weighted by the number of messages at that point.
  - Parameters set: 휀휀 = 2,000 meters and 푁푁 = 1,000.
- Interpretation of 푁푁:
  - Calibration rationale: the minimum number of observations 푁푁 is chosen so a single-berth terminal must be occupied at least about five percent of the time to be identified as a cluster.
  - With 12 daily messages over four years, a constantly occupied berth would have 12x365x4 = 17,520 messages; DBSCAN parameter set to 1,000.
- Sensitivity:
  - Figure 3 (described) shows sensitivity in terms of number of core and noise points under various choices of 푁푁; by and large, number of noise points does not change dramatically under different choices of 휀휀.
- Distance computations:
  - Distances on the map (including between message locations, and also between centroids and ports) were computed using the haversine formula.

### Assigning AIS-derived polygons to countries and ports
- Centroid-to-country assignment:
  - Use version 10 of the Maritime Boundaries Geodatabase (Flanders Marine Institute, 2018) to map centroids to sovereign countries via territorial waters polygons.
  - If a centroid search over water returns no sovereign country (sometimes centroids fall over land), use polygons over land areas from the World Borders Dataset.
  - Outcomes: Of around 3,500 cluster centroids:
    - All but around 1,300 are mapped to a country based on maritime boundaries.
    - Of those 1,300, all but around 500 are mapped to a country based on land boundaries.
    - Final result: about 3,000 centroids mapped to a country and about 500 centroids not assigned to any country.
- Assigning centroids to head ports (hierarchical two-step procedure):
  - Step 1: If centroid has a country assigned, search for head ports within a 30 km radius that belong to the same country. If non-empty, assign polygon to the head port in this set closest to the centroid.
  - Step 2: If no head port in the same country within 30 km, assign polygon to the WPI port in the same country that is closest to the centroid.
  - If centroid does not have a country assigned: search for head ports within a 30 km radius; if non-empty assign to closest head port in that set; otherwise assign to the WPI port that is closest to the centroid.
  - Rationale: Hierarchical procedure avoids assigning a cluster to a non-head port d belonging to head port j when head port k is nearer; 30-km calibration consistent with port-group construction.
- Removing distant assignments and synthetic ports:
  - If centroid-to-nearest-port distance > 75 km, the assignment is removed.
  - Clusters removed by this rule are termed "synthetic ports".
  - Synthetic ports that lie on the waters or land of a country are synthetic ports of that country.
  - Total synthetic ports: 16, which are assigned to the closest countries and manually validated to minimize misclassification.

### Identifying port calls – supervised learning overview
- Problem statement:
  - Not all messages from within a port polygon indicate true port calls (ships may traverse port polygons without stopping to load/unload).
  - Examples: ships transiting river channels (Rotterdam, Buenos Aires), crossing San Francisco Bay to reach Oakland, passing through Singapore Strait.
- Reference data for supervised labeling:
  - Use official vessel entrances and clearances at U.S. ports from the U.S. Army Corps of Engineers’ Navigation Data Center (NDC), compiled in partnership with U.S. Department of Homeland Security’s Customs and Border Protection agency (national waterway data).
  - NDC data record most port calls on U.S. ports and include vessels’ IMOs and date of entry.
  - As of March 2020, NDC data are available through end 2018.
- Treatment of U.S.-flagged vessels:
  - By regulation, all vessels except U.S.-flagged ships coming directly from another U.S. port and without foreign goods onboard must file an entrance statement.
  - U.S.-flagged ships make up only 0.5 percent of the dataset.
  - Analyses in this section restrict attention to non-U.S.-flagged vessels that in the raw database arrive from non-U.S. ports.
- Matching rules between AIS-derived polygon visits and NDC:
  - Matched port calls:
    - For each NDC port call, search for same IMO in port visit data over a window of +/- 2 days around the NDC-recorded date.
    - Rationale for +/- 2 days:
      - By law vessels have up to two days to file their entrance report.
      - UTC vs local U.S. time discrepancies (timestamps in AIS are UTC; NDC records local U.S. date): shift AIS timestamps by six hours across the board to limit but not eliminate this discrepancy.
      - Vessels may report entry ahead of entering the port (e.g., while waiting at anchorage).
    - When multiple port calls within the window exist, match NDC records to the ones with closest dates; if tie persists, prefer the earliest record.
  - False negatives:
    - IMOs recording port calls in NDC for which no port call is found in AIS-derived data over the +/- 2-day window.
  - False positives:
    - IMOs recording port calls in AIS-derived data that remain unmatched to NDC port calls over the +/- 2-day window.
    - Note: some vessels are not required to file an entrance record, so a false positive does not necessarily imply the vessel did not enter a port.

### Mapping results and assessing blind spots (2015–2018)
- Aggregated counts and matching summary:
  - For period 2015-2018:
    - AIS-derived GTI polygons record more than 500,000 distinct ship-date pairs with messages coming from within polygons associated with U.S. ports.
    - NDC shows nearly 290,000 entrance records.
    - Around 215,000 NDC port calls can be mapped to a GTI polygon visit within a +/- 2-day window.
  - Initial classification from matching:
    - Matched port calls: frequency distribution (from Table 1):
      - Found by GTI on day t = -2: 80; FrequencyPercentCumulative line shows "-28040.10.1" (table reproduction preserved as in source).
      - Found by GTI on day t = -1: 9,458 (1.6 percent; cumulative 1.7 percent).
      - Found by GTI on day t = 0: 189,538 (31.4 percent; cumulative 33.1 percent).
      - Found by GTI on day t = +1: 104,427 (17.3 percent; cumulative 34.8 percent).
      - Found by GTI on day t = +2: 4,460 (0.7 percent; cumulative 35.5 percent).
    - In ML-AIS but not in NDC (false positives): 318,168 (52.6 percent; cumulative 88.1 percent).
    - In NDC but not in ML-AIS (false negatives): 71,674 (11.9 percent; cumulative 100.0 percent).
    - Total observations: 604,529 (100.0 percent).
- Interpretation of false negatives and dataset coverage:
  - Many of the apparent false negatives correspond to IMOs that never appear in the AIS-based data (mostly small ships such as tugs, barges, service vessels that do not participate in international trade).
  - Of the around 71,000 false negative port calls, about 58,000 correspond to IMOs not in the AIS data.
  - Excluding these IMOs from the analysis reveals higher matching performance (text indicates subsequent analysis shows algorithm can match but detailed numbers beyond this exclusion are in the following text not included here).

_Italic: Source — wpiea2020057-print-pdf (canonical PDF content provided) _

### 94.3 percent of NDC’s port calls (Table 2).

### 94.3 percent of NDC’s port calls (Table 2)

### Detection performance of AIS-derived port polygons
- At 94.3 percent, the GTI polygons capture the overwhelming majority of official U.S. port calls.
- Total labeled official port calls used in some tables: Total = 227,621 (100.0).
- In NDC but not in ML-AIS (false negatives): 12,934 (5.71 percent of the total).
- Example by vessel type (from Table 3):
  - CONTAINER SHIP: Total official port calls = 71,598; Found by GTI = 69,599; False negative rate = 2.79 percent.
  - BULK CARRIER: Total = 34,882; Found = 31,666; False negative rate = 9.22 percent.
  - OIL/CHEMICAL TANKER: Total = 32,324; Found = 30,375; False negative rate = 6.03 percent.
  - GENERAL CARGO: Total = 20,490; Found = 19,232; False negative rate = 6.14 percent.
  - VEHICLES CARRIER: Total = 18,965; Found = 18,347; False negative rate = 3.26 percent.
  - CRUDE OIL TANKER: Total = 18,182; Found = 17,352; False negative rate = 4.56 percent.
  - SELF DISCHARGING BULK CARRIER: Total = 6,608; Found = 5,422; False negative rate = 17.95 percent.
  - LPG TANKER: Total = 5,993; Found = 5,683; False negative rate = 5.17 percent.
  - OIL PRODUCTS TANKER: Total = 4,185; Found = 3,877; False negative rate = 7.36 percent.
  - RO-RO CARGO: Total = 3,348; Found = 3,248; False negative rate = 2.99 percent.
- False negative rates are below 10 percent for all presented vessel types except self-discharging bulk carriers; smaller vessels within that class are more likely missed (Appendix A, Figure A1).
- False negative rates vary by port geography; analysis by U.S. port and vessel type was performed (example illustrated for bulk carriers in Figure 5).

### False positives and supervised learning approach
- A pervasive issue: many AIS-derived polygon visits step on polygons but lack corresponding NDC entry records (false positives).
- Observed patterns motivating filtering:
  - For GTI port visits with matching NDC port call, about half have a previous port visit in the U.S.; this fraction rises to nearly 70 percent for GTI visits without a matching NDC port call.
  - In GTI port visits with matching NDC port call, about 3 percent of vessels are U.S.-flagged; this fraction rises to 17 percent for GTI visits without a matching NDC port call.
- To reduce bias, observations where the previous GTI port of call is in the U.S. and the vessel is U.S.-flagged are dropped from classifier training.
- Intuition features used to distinguish true port calls from false positives: average message speed, duration within polygon, number of anchored/moored messages, etc. Fast transits (e.g., average speed ~10 knots) and durations less than one hour are unlikely to represent cargo operations.

### Random forest classifier: features, training, and performance
- Motivation: non-linear and non-monotonic relations between covariates (e.g., average speed, duration) and port-call probability call for flexible, nonparametric models.
- Classifier type: Random forest ensemble, using Gini's diversity index as cost function and deadweight-tonnage-weighted frequencies to penalize misclassification proportionally to vessel gross tonnage.
- Features: 16 covariates (Table 4), including:
  - Anchored message ratio (Fraction of messages with status anchored over anchored+moored)
  - Total no. anchored/moored messages
  - Is anchorage polygon (1 if anchorage polygon)
  - Hours from draught change
  - Dummy previous mooring call (1 if other mooring call in past 48 hours)
  - Hours from previous mooring call
  - No. draught decrease / No. draught increase
  - Share of low-speed messages (speed < 0.5 knots)
  - No. messages
  - Average message speed
  - Average speed (distance over time from first message out to first message in polygon)
  - Duration of stay (time from first message out to first message in polygon)
  - Summer deadweight tonnage
  - Gross tonnage
  - Ship type
- Draught considerations:
  - Draught is manually entered and prone to measurement error and lags.
  - Use last observed draught before entering polygon, presuming more careful reporting before port entry.
- Training and tuning:
  - Labeled port visits: 472,085.
  - Test set: randomly pick 10 percent; remaining 90 percent used for training and hyperparameter tuning with five-fold cross validation.
  - Hyperparameter tuning via Bayesian optimization.
  - Tuned hyperparameters: minimum number of points at each leaf node = 1; number of trees = 232.
  - Randomness: bootstrap fraction 0.6 of training samples for each tree.
- Model comparison on testing sample (Table 5):
  - Accuracy: RANDOM FORESTS 91.3%, SVM 84.5%, LOGISTIC REGRESSION 84.7%.
  - Precision: RANDOM FORESTS 89.4%, SVM 80.9%, LOGISTIC REGRESSION 81.2%.
  - Recall: RANDOM FORESTS 91.1%, SVM 85.0%, LOGISTIC REGRESSION 84.8%.
- Interpretation:
  - Random forest outperforms SVM and logistic regression on accuracy, precision, and recall, better handling non-linearities.
  - Precision (near 90%) indicates low false positive rate relative to alternatives; recall (91.1%) indicates low false negative rate.
  - ROC curves (Figure 7) show random forest as the best method among the three.
- Feature importance (Figure 8):
  - Top two features: anchored message ratio and duration of stay within polygon.
  - Other important predictors: number of messages from polygon, ship type, time to next mooring polygon.

### From port calls to voyages and volume of trade
- After classifying polygon visits as port calls (1) or not (0) using trained random forest, the classifier is applied to worldwide AIS-derived polygon-visit dataset to produce estimated global port call database.
- Voyages constructed between estimated port calls; certain ports are “seen through” and discarded to avoid misclassification of bilateral origin/destination:
  - Jurisdictions discarded: Panama and Gibraltar; two Suez Canal ports (As Suways, Bur Said (Port Said)).
- Metric tons estimation per inbound ship visit:
  - Notation:
    - m_{dmit}^{IN} = inbound metric tons of cargo transported by ship d at time t into a port.
    - dwt_i = deadweight tonnage of ship i.
    - d_{i,mdes} = maximum (design) draught of ship i.
    - d_{i,ballast} = draught when not transporting any goods (ballast draught).
    - d_{iit} = observed draught.
  - Estimation formula:
    - m_{dmit}^{IN} = dwt_i × (d_{iit} − d_{i,ballast}) / (d_{i,mdes} − d_{i,ballast})
    - (Adjusts ship’s total capacity with current utilization rate.)
- Ballast draught imputation:
  - For each vessel, compute ratio: first percentile of observed draught over four years divided by design draught.
  - By ship type and deadweight tonnage tertile, take median of this ratio, denote mmr_{rbd}.
  - Impute vessel ballast draught as: d_{i(r,b),ballast} = mmr_{rbd} × d_{i,mdes}.
- Mapping metric tons to imports/exports/internal trade (procedure):
  1. Define outgoing cargo weight for ship visit as m_{dmit}^{OUT} = m_{dmi,t+1}^{IN} (outgoing = incoming weight at next port). Net cargo offloaded at port: m_{dmit}^{OFF} = m_{dmit}^{IN} − m_{dmit}^{OUT} (can be negative or positive).
  2. If ship enters a port in a new country (previous port in different country) and m_{dmit}^{OFF} > 0 → label as import; if m_{dmit}^{OFF} < 0 → do not count imports on this trip (may be internal trade).
  3. If first port of call in country had an import event, count subsequent drops in metric tons as imports until the first instance the ship starts gaining metric tons (i.e., until m_{dmi,t+(J+1)}^{OUT} < 0). Exports defined analogously by scanning backwards from last port call in the country and labeling contiguous gains as exports.
- Aggregation:
  - Group port-call events into importing, exporting, or internal trade and aggregate metric tons by desired country grouping, vessel type, and time period (e.g., monthly country-level import/export volume indicators, weekly bilateral crude oil trade).

### Results and benchmarking
- Macroeconomic nowcasting:
  - Global seasonally adjusted monthly import volume gauge correlation with official statistics (CPB): 0.88.
  - In 3m/3m growth rate terms, global correlation = 0.4.
  - World exports show similar performance; growth rate of exports somewhat weaker.
  - Median correlation across CPB sample (3-month moving averages in levels): imports = 0.68, exports = 0.46.
  - Median correlations in growth rates: imports = 0.27, exports = 0.18.
  - Example: Argentina exports show high correlation with GTI (Table 6, Figure 10).
  - Benchmarking limited to 43 economies (including Euro Area) plus the world due to scarcity of monthly trade volume data.
- Sectoral analysis — crude oil:
  - Comparison with JODI monthly crude oil trade (thousands of metric tons).
  - Global crude oil imports:
    - Correlation in raw levels: 0.73.
    - Correlation on 3-month moving average basis: 0.81.
    - Correlation in 3m/3m growth rates: 0.47.
  - Exports: tracking poorer at global level; JODI export data quality and reporting differences likely contributors (JODI global imports exceed exports by over 15 percent; assessment-code quality differences noted).
  - Country-level mismatches can arise where port of arrival differs from final destination (e.g., Germany, Netherlands).
- Event studies and daily-frequency applications:
  - Hurricane Maria (Puerto Rico, landfall September 20, 2017): offloading vessel traffic declined by as much as 75 percent and took around 10 days to recover.
  - COVID-19 pandemic examples:
    - Port-level AIS message frequency visuals for Port of Ningbo across weeks in 2020 show dramatic reductions during strict containment and partial rebounds by April.
    - Daily estimated metric tons of exports for China (normalized to 2017–2019 avg, Lunar New Year-adjusted) show dramatic fall with resumption in early–mid March but incomplete recovery by mid-April.
  - Real-time world trade estimates (30-day moving averages, relative to 2017–2019 average) show sectoral heterogeneity: oil-related exports strong at times (consistent with storage at sea) while container and vehicle exports/imports show weak readings during production halts and demand drops.

### Ongoing developments and avenues for improvement
- Algorithmic refinements not requiring new data:
  - Upgrade port-polygon construction: DBSCAN limitations (constant density threshold) motivate testing OPTICS (Ankerst et al., 1999) to find contours where ship density drops and better separate low-traffic ports from nearby high-traffic ports.
  - Predict ship destination while at sea using unstructured AIS textual info and voyage history to increase forecasting horizon from ~1 month to 2–3 months.
- Improvements requiring additional data:
  - More detailed vessel information (ship registers) to refine cargo weight estimation:
    - Current method assumes linear change in cargo weight with draught (equivalent to rectangular cuboid hull); more accurate hull form coefficients (e.g., block coefficient) would improve metric-ton calculations (Jia, Prakash, and Smith, 2019; MAN, mimeo).
    - Direct ballast draught values from registers would be preferable to imputed values.
- Domain-specific modifications:
  - Ports with unusual geographies (oil platforms, FPSOs, riverine ports) and special operations (lightering) can cause double counting or missed events; expert knowledge and port-specific rules needed for further gains.
- Trade-weighted indices:
  - Alternative approach mapping coarse vessel classification to HS 4-digit codes and applying COMTRADE-based weights was tested (Appendix B, Table A1).
  - Results: some economies improve, others do not; imports in levels show more unambiguous improvement than growth rates. Not a clear overall improvement in isolation.

### Concluding remarks
- The methodology provides an end-to-end proof of concept to construct trade volume indicators relying only on AIS data and off-the-shelf machine learning.
- Achieves good fit with official trade statistics for many countries and the world in aggregate; useful for sectoral crude oil analysis and granular event studies.
- Further improvements expected from algorithmic refinements, additional vessel data, and country-specific domain knowledge.
- Expanding availability of monthly official trade volume statistics for more countries would aid further benchmarking and application to emerging and developing economies that rely heavily on seaborne trade.

*Source: wpiea2020057-print-pdf - 94.3 percent of NDC’s port calls (Table 2).*

### REFERENCES

### REFERENCES

### Machine learning, clustering, and statistical methods
- Ankerts, M., M.M. Breunig, H.-P. Kriegel and J. Sander. 1999. “OPTICS: Ordering Points To Identify the Clustering Structure,” Proceedings ACM SIGMOD, International Conference on Management of Data.
- Breiman, L. 2001. “Random forests,” Machine learning 5-32.
- Cortes, C. and V. Vapnik. 1995. “Support-Vector Networks,” Machine Learning, 20, 273-297.
- Ester, M., H.-P. Kriegel, J. Sander, X. Xu. 1996. “A density-based algorithm for discovering clusters in large spatial databases with noise,” Proceedings of the Second International Conference on Knowledge Discovery and Data Mining.
- Friedman, J., T.  Hastie, and R. Tibshirani. 2001. The elements of statistical learning. New York: Springer series in statistics.
- Forgy, E.W. 1965. “Cluster analysis of multivariate data: efficiency versus interpretability of classifications,” Biometrics. 21 (3): 768–769.
- Getis, A. 1992. “The Analysis of Spatial Association by Use of Distance Statistics,” Geographical Analysis, Vol. 24, No. 3.
- Ho, T.K. 1995. “Random Decision Forests,” the 3rd International Conference on Document Analysis and Recognition. Montreal, QC: IEEE. 278–282.
- Ho, T.K. 1998. “The random subspace method for constructing decision forests,” IEEE transactions on pattern analysis and machine intelligence 832-844.
- Kishore, N., D. Marques, A. Mahmud, M.V. Kiang, Ii Rodriguez, A. Fuller, P. Ebner, C. Sorensen, F. Racy, J. Lemery, L. Maas, J. Leaning, et al. 2018. “Mortality in Puerto Rico after Hurricane Maria,” New England Journal of Medicine, 379, 2.
- Kohavi, Ron. 1995. “A study of cross-validation and bootstrap for accuracy estimation and model selection,” the International Joint Conference on Artificial Intelligence.
- Lloyd, S. 1982. “Least squares quantization in PCM,” IEEE Transactions on Information Theory, 28 (2): 129–137.
- Snoek, J., H. Larochelle, R. P. Adams. 2012. “Practical Bayesian optimization of machine learning algorithms,” Advances in neural information processing systems 2951-2959.
- Spiliopoulos G., D. Zissis and K. Chatzikokolakis. 2018. “A Big Data Driven Approach to Extracting Global Trade Patterns,” in: Doulkeridis C., Vouros G., Qu Q., Wang S. (eds) Mobility Analytics for Spatio-Temporal and Social Data. MATES 2017. Lecture Notes in Computer Science, vol 10731.

### AIS, vessel traffic, shipping, and maritime data applications
- Adland, R., H. Jia and S.P. Strandenes. 2017. “Are AIS-based trade volume estimates reliable? The case of crude oil exports,” Maritime Policy & Management, Vol. 44, No. 5, pp. 657-665.
- Arslanalp, S., M. Marini and P. Tumbarello. 2019. “Big Data on Vessel Traffic: Nowcasting Trade Flows in Real Time,” IMF Working Paper 19/275.
- Hammer, C.L., D.C. Kostroch and G. Quirós. 2017. “Big Data: Potential, Challenges, and Statistical Implications,” IMF Staff Discussion Note 17/06.
- Heiland, I., A. Moxnes, K.H. Ulltveit-Moe and Y. Zi. 2019. “Trade From Space: Shipping Networks and The Global Implications of Local Shocks,” mimeo.
- Jia, H., V. Prakash, T. Smith. 2019. “Estimating vessel payloads in bulk shipping using AIS data,” Int. J. Shipping and Transport Logistics, Vol. 11, No. 1, 2019 25
- Natale, F., M. Gibin, A. Alessandrini, M. Vespe and A. Paulrud. 2015. “Mapping fishing effort through AIS data,” PloS one, 10(6).
- Liu, H., Z. Meng, Z. Lv et al. 2019. “Emissions and health impacts from global shipping embodied in US–China bilateral trade,” Nat Sustain 2, 1027–1033.
- MAN. Mimeo. “Basic Principles of Ship Propulsion,” available at https://spain.mandieselturbo.com/docs/librariesprovider10/sistemas-propulsivos-marinos/basic-principles-of-ship-propulsion.pdf
- National Geospatial-Intelligence Agency (NGA). 2017. World Port Index 2017.
- Flanders Marine Institute (2018). Maritime Boundaries Geodatabase: Maritime Boundaries and Exclusive Economic Zones (200NM), version 10.
- Spiliopoulos G., D. Zissis and K. Chatzikokolakis. 2018. “A Big Data Driven Approach to Extracting Global Trade Patterns,” in: Doulkeridis C., Vouros G., Qu Q., Wang S. (eds) Mobility Analytics for Spatio-Temporal and Social Data. MATES 2017. Lecture Notes in Computer Science, vol 10731.
- Spiliopoulos entry also appears under Machine learning, clustering, and statistical methods.

### International finance, macroeconomics, and crises
- Catao, L. and G.M. Milesi-Ferretti. 2014. “External liabilities and crises,” Journal of International Economics, Vol. 94 (1).
- Frankel, J.A. and A.K. Rose. 1996. “Currency Crashes in Emerging Markets: An Empirical Treatment,” International Finance Discussion Papers, Federal Reserve Board of Governors.
- Kaminsky, G., S. Lizondo and C.M. Reinhart. 1998. “Leading Indicators of Currency Crises,” IMF Staff Papers, Vol. 45 (1).
- Gopinath, G. 2020. “The Great Lockdown: Worst Economic Downturn Since the Great Depression,” IMF Blog, April 14, 2020.

### Trade, maritime transport, and related reviews
- Brancaccio, G., M. Kalouptsidi and T. Papageorgiou. Forthcoming. “Geography, Transportation and Endogenous Trade Costs,” Econometrica.
- United Nations Conference on Trade and Development (UNCTAD). 2017. Review of Maritime Transport.

### Other data, surveys, and domain-specific sources
- Smith, A.B. 2020. “2010-2019: A landmark decade of U.S. billion-dollar weather and climate disasters,” National Oceanic and Atmospheric Administration.
- MAN. Mimeo. “Basic Principles of Ship Propulsion,” available at https://spain.mandieselturbo.com/docs/librariesprovider10/sistemas-propulsivos-marinos/basic-principles-of-ship-propulsion.pdf

*Source: REFERENCES (wpiea2020057-print-pdf)*

---


_Source: https://www.imf.org/-/media/files/publications/wp/2020/english/wpiea2020057-print-pdf.pdf_
