| Grooming rule | Description | Fields affected | Time spent |
|---|---|---|---|
| FLKIN | Update estimated catch species to SUR when KIN is reported from diving events with no MHR support; Update target species to SUR when KIN is reported from diving events with no MHR support; Update landed species to SUR when KIN is landed from trips with diving events and no MHR support | 376 | 76.8 secs |
| ESTGT | Create estimated catch records for events with a total catch weight only | 46216 | 80.3 secs |
| LADAM | Landings where the landing date is missing | 0 | 32.9 secs |
| LADAF | Landings where the landing date is in the future | 0 | 30.6 secs |
| LADTI | Invalid landing destination | 8547 | 39.7 secs |
| LAFLA | Correct landings using a flatfish species code to FLA | 410651 | 1281.0 secs |
| LAHPB | Correct landings using a groper species code to HPB | 112606 | 53.3 secs |
| LAOEO | Correct landings using an oreo species code to OEO | 41295 | 41.5 secs |
| LASQU | Recode SQU1J and SQU1T landings to SQU1 | 187582 | 15.9 secs |
| LATUN | Correct stock code for non-QMS tunas | 9655 | 2.1 secs |
| LASEC | Landings to Crown or experimental stock codes | 17232 | 46.1 secs |
| LAQMS | Replace pre-QMS pseudo-stock with the post-QMS stock code | 116778 | 133.7 secs |
| LADMR | Mandatory returns (e.g. sub-MLS) | 359333 | 36.7 secs |
| LADTH | Retained (non-final) landings | 1077067 | 99.2 secs |
| LADTT | Vessel received transhipments | 62570 | 56.0 secs |
| LASCF | Correct some state codes | 3926 | 11.8 secs |
| LASCI | Landings to invalid state code | 21414 | 12.1 secs |
| LASCD | Drop landings of secondary product states | 162667 | 99.7 secs |
| LADUP | Duplicate landings | 88366 | 86.0 secs |
| LACFM | Replace missing conversion factors with the median over all years | 2813557 | 453.5 secs |
| LAGWI | Estimate missing greenweights | 397572 | 112.0 secs |
| LAGWM | Missing greenweights that cannot be estimated | 43225 | 44.7 secs |
| LAGWO | Identify and fix order of magnitude errors in landings | 103724 | 883.1 secs |
| DCFxx | NA | 0 | 190.7 secs |
| FEMDV | Update historical diving method codes to DV | 176779 | 26.1 secs |
| FEPMN | Add PSH as a method code for certain vessels if method is null | 164 | 9.5 secs |
| FEPMI | Replace missing methods if there is only one method used on the trip (by form type) | 975 | 39.8 secs |
| FEPMM | Flag trips if any events have a missing method | 9417 | 8.9 secs |
| FESAI | Substitute the modal statistical area from a trip for missing areas | 40894 | 14.6 secs |
| FESAM | Flag events with missing statistical areas | 0 | 3.5 secs |
| FESAS | For BCO 4 only correct RL statistical areas to general areas | 10630 | 332.4 secs |
| FESAF | Flag non RLP events using RL statistical area codes | 13607 | 10.4 secs |
| FESDF | Flag events in the future | 0 | 4.2 secs |
| FESDM | Flag events with missing start date/time | 0 | 3.7 secs |
| FETSE | Set target species to group code for FLA, HPB and OEO species | 215529 | 10.5 secs |
| FETSW | Flag and set target species to null if target species is not a valid species code | 146888 | 18.8 secs |
| FETSI | Replace missing target species with the modal value for a trip | 4510 | 10.3 secs |
| FEETN | Flag and fix some CP effort errors | 0 | 25.9 secs |
| FEEHN | Fix transposed effort numbers for lining methods on CELR forms | 14000 | 18.8 secs |
| FEEMU | Fix SN mesh sizes recorded in inches | 25012 | 9.8 secs |
| FEFMA | Mark trips which landed to more than one fishstock for straddling statistical areas | 0 | 1.0 secs |
| FEMEM | Flag events where the primary effort measure is missing | 9978 | 11.5 secs |
| FEHDE | Flag records where the maximum daily effort is out of range | 4819 | 40.1 secs |
| FEDBE | Transpose bottom and effort depths if reported effort depth > bottom depth | 220994 | 15.6 secs |
| ESCWN | Correct cases where estimated catch is recorded in weight but number of fish is expected | 22904 | 336.0 secs |
| PRSCI | Processed catch with invalid state code | 1428383 | 104.6 secs |
| PRSCD | Drop processed catch of secondary product states | 370967 | 24.1 secs |
| PRCFM | Replace missing conversion factors with the median over all years | 512072 | 32.5 secs |
2 Grooming
The term grooming refers to flagging or fixing errors in the source data. Grooming is undertaken when the kahawai database is built. At present, grooming is primarily implemented for catch, effort, and landings data in the kahawai_edw schema.
2.1 Grooming of catch, effort, and landings data
Fishing effort undertaken by commercial fishers in New Zealand, the catches estimated for each fishing event, and the landings and disposals of catch from every commercial fishing trip are recorded in MPI’s Enterprise Data Warehouse (EDW) and managed in the kahawai_edw schema in the kahawai database.
Grooming of these data is necessary for two key reasons:
- the data are provided by humans and some errors are inevitable; and
- to support fisheries compliance and enforcement activity, MPI has typically stored data “as reported” by fishers, even when errors are identified.
Prior to the introduction of the Electronic Reporting System (ERS) in the late 2010s, fishers provided data to MPI using paper forms and data were transcribed into the database. Common errors on paper forms include:
- the wrong value being entered; this includes order of magnitude errors due to misplaced decimal points, and “180 errors” caused by using the wrong suffix (E or W) for longitudes near the 180th meridian;
- values being entered in the wrong field; this was a particular problem for reporting that used the Catch, Effort, and Landing Return (CELR), a multipurpose form where the use of the fields differed depending on the fishing method;
- misinterpreted values due to hard to read handwriting or messy forms; and
- transcription errors made when forms were entered into the database.
The WAREHOU database (Ministry of Fisheries (2010)) provided some opportunity for data corrections to be made by supporting three versions of a form:
- the literal version containing the latest version of the information that a fisher provided to the Ministry;
- the interpreted version of a form where at least one field of data had been changed from the literal version. For example, if a fisher wrote “snapper” this could be interpreted as the species code “SNA”. Permissible interpretations of the data were formalised and documented when data entry was delegated to FishServe;
- provision for a research version of a form was made, but not implemented, because scientists may disagree on the correct interpretation, or may need to interpret the data in different ways for different analyses.
Data extracts loaded into the kahawai database use the interpreted version of these data, where available. The grooming process implemented in the kahawai database is intended to provide a more flexible approach than that envisaged in the WAREHOU database. The default grooming applied in the kahawai database draws on experience accumulated over analyses of many different stocks and fisheries, undertaken by MPI’s fisheries science contractors, and peer reviewed by the working group process. Furthermore, use of any given grooming rule is optional; the raw data are available, and the stepwise changes made by particular rules can be reversed when the data are analysed.
The introduction of electronic reporting has reduced the opportunity for making some of the errors that were easily made on paper forms. For example, positions can usually be entered automatically into the reporting application directly from a GPS. Likewise, only valid species and method codes can be entered. However, errors are still possible: the correct values must be chosen, and conveniences in the reporting applications (such as default values maintained between events) may also lead to undetected errors.
2.1.1 Grooming applied in the kahawai build
The grooming rules applied to catch, effort, and landings data in the latest kahawai database build are listed in Table 2.1. The naming of the grooming rules is based on the convention implemented by Bentley (2012) who, in turn, was aiming to standardise and extend the grooming approach documented by Starr (2007).
Most of the grooming rules implemented by Bentley (2012) were reimplemented in the kahawai database in the mid-2010s and have subsequently been updated and expanded. While Bentley (2012) envisaged applying the rules to project-specific data extracts, the extensive extracts used in the kahawai database mean that the grooming code, and also the allocation procedures, have the full dataset available and so should be best positioned to identify abberant data.