scieee Open visual document viewer

Deep Learning based Detection, Segmentation and Counting of Benthic Megafauna in Unconstrained Underwater Environments

Lütjens, Mona Caroline,Sternberg, Harald

Abstract

Assessing and monitoring benthic communities is increasingly important in view of global alteration of marine environments. Deep learning has proven to effectively detect marine specimen in underwater imagery but still face problems with small input datasets, unconstrained environments and class imbalance. This study evaluates a data augmentation strategy to alleviate these limitations. Through synthetically derived image compositions, the entire input dataset was greatly extended from 700 to 12700 images. Additionally, specimen numbers of brittle stars, soft corals and glass sponges are equalized resulting in a mean average precision increase of 24 %. The overall mean average precision for box detections yields 76.7 and for instance segmentation 67.7 at an intersection over union threshold of 0.5. This study shows that deep architectures such as the deployed CenterMask via ResNeXt-101 model can successfully be trained with few original images from varying underwater scenes.

Full text

IFAC Pape sOnLine 54-16 (2021) 76–82 ScienceDi ec A ailable online a www.sciencedi ec .com 2405-8963 Copy igh © 2021 The Au ho s. This is an open access a icle unde he CC BY-NC-ND license . Pee e iew unde esponsibili y o In e na ional Fede a ion o Au oma ic Con ol. 10.1016/j.i acol.2021.10.076 10.1016/j.i acol.2021.10.076 2405-8963 Copy igh © 2021 The Au ho s. This is an open access a icle unde he CC BY-NC-ND license ( h ps://c ea i ecommons.o g/licenses/by-nc-nd/4.0/ ) Deep Lea ning based De ec ion, Segmen a ion and Coun ing o Ben hic Mega auna in Uncons ained Unde wa e En i onmen s Mona Lü jens* Ha ald S e nbe g* *Ha enCi y Uni e si y Hambu g, Depa men o Hyd og aphy and Geodesy, Henning-Vosche au-Pla z 1, 20457 Hambu g, Ge many (e-mail: {mona.lue jens,ha ald.s e nbe g}@hcu-hambu g.de) Abs ac : Assessing and moni o ing ben hic communi ies is inc easingly impo an in iew o global al e a ion o ma ine en i onmen s. Deep lea ning has p o en o e ec i ely de ec ma ine specimen in unde wa e image y bu s ill ace p oblems wi h small inpu da ase s, uncons ained en i onmen s and class imbalance. This s udy e alua es a da a augmen a ion s a egy o alle ia e hese limi a ions. Th ough syn he ically de i ed image composi ions, he en i e inpu da ase was g ea ly ex ended om 700 o 12700 images. Addi ionally, specimen numbe s o b i le s a s, so co als and glass sponges a e equalized esul ing in a mean a e age p ecision inc ease o 24 %. The o e all mean a e age p ecision o box de ec ions yields 76.7 and o ins ance segmen a ion 67.7 a an in e sec ion o e union h eshold o 0.5. This s udy shows ha deep a chi ec u es such as he deployed Cen e Mask ia ResNeX -101 model can success ully be ained wi h ew o iginal images om a ying unde wa e scenes. Keywo ds: objec de ec ion, deep lea ning, da a augmen a ion, ma ine image y, ben hic mega auna 1. INTRODUCTION Global al e a ion o ma ine en i onmen s due o o e ishing, pollu ion, physical habi a des uc ion and clima e change ha e led o an inc easing decline o animal di e si y and abundance h oughou many ma ine ecosys ems (Jackson, 2008). Especially ben hic mega auna in he Sou he n Ocean a e a isk o en i onmen al change and a e o signi ican ecological alue as hey al e small-scale opog aphy o seabed habi a s a ec ing he en i e ben hic communi y (Gili e al., 2006). Assessing he biodi e si y and cha ac e isa ion o ben hic communi ies is inc easingly impo an o iden i ying ulne able ma ine ecosys ems and de eloping conse a ion s a egies. Ma ine habi a s ha e been s udied based on mainly h ee p ac ices: physical sampling me hods using sledges o awling, acous ical echniques and op ical sys ems. While physical me hods a e able o collec specimen on a lowe axonomical scale hey des uc he en i onmen and hei sampling a e is a he low. Op ical sys ems a e mo e cos e ec i e, obus and p ecise han acous ic sys ems which led o a la ge g owing lib a y o digi al unde wa e images h oughou ecen yea s pa ing he way o new esea ch in au oma ic analy ical me hods. Objec de ec ion, classi ica ion and segmen a ion o image y has become a subs an ial ask in he ield o compu e ision. P io adi ional ea u e desc ip o s a e o en used o ex ac colou , shape o ex u e in o ma ion and a e good a de ec ing speci ic single objec s such as scallops (Dawkins e al., 2013) o lobs e s using classi ie s such as suppo ec o machine (Tan e al., 2018). Howe e , hey a e no obus o a ying unde wa e scenes ha a e exposed o ma ine snow, wa e u bidi y, lens dis o ion, spa se, uns able illumina ion and colou shi due o he su ey pla o m’s a ia ion in speed, angle o al i ude. Mo eo e , he e is a high a iabili y and in a iabili y be ween ea u es belonging o he same class and di e en classes, espec i ely (Pa oni e al., 2021). Deep lea ning (DL) using con olu ional neu al ne wo ks ha e p o en o ou pe o m adi ional based objec de ec ion (Gonzalez-Cid e al., 2017) as hey a e mo e in a ian o he de o ma ion o images. Addi ionally, images can di ec ly be used as inpu wi hou he necessi y o p e-p ocessing. Be e esul s a e o en achie ed using deepe laye s hus inc easing he numbe o pa ame e s o se e al millions. To achie e high pe o mances wi h hose models using images o uncons ained unde wa e scenes o ac oss a ying pla o ms, a la ge aining da ase size is he mos c ucial pa (Langenkämpe e al., 2020), o en, howe e , e y ime- consuming and cos ly o es ablish. This pape in es iga es he e ec o inpu image da a augmen a ion and composi ion s a egies in an a emp o o e come limi a ions o la ge da a se size and class imbalance. Taking accoun o he esul s, he s a e-o - he a ancho - ee objec de ec ion and ins ance segmen a ion model Cen e Mask (Lee and Pa k, 2019) ia ResNeX -101 (Xie e al., 2017) will be ained based on a small, highly di e se 700 image da ase . In his wo k, he de ec ion, segmen a ion and coun ing o glass sponges (hexac inellids), so co als (p imnoids and ch ysogo giids) and b i le s a s (ophiu oids) will be assessed p o iding i s s eps owa ds u u e abundance and size es ima ions o hese specimen. Deep Lea ning based De ec ion, Segmen a ion and Coun ing o Ben hic Mega auna in Uncons ained Unde wa e En i onmen s Mona Lü jens* Ha ald S e nbe g* *Ha enCi y Uni e si y Hambu g, Depa men o Hyd og aphy and Geodesy, Henning-Vosche au-Pla z 1, 20457 Hambu g, Ge many (e-mail: {mona.lue jens,ha ald.s e nbe g}@hcu-hambu g.de) Abs ac : Assessing and moni o ing ben hic communi ies is inc easingly impo an in iew o global al e a ion o ma ine en i onmen s. Deep lea ning has p o en o e ec i ely de ec ma ine specimen in unde wa e image y bu s ill ace p oblems wi h small inpu da ase s, uncons ained en i onmen s and class imbalance. This s udy e alua es a da a augmen a ion s a egy o alle ia e hese limi a ions. Th ough syn he ically de i ed image composi ions, he en i e inpu da ase was g ea ly ex ended om 700 o 12700 images. Addi ionally, specimen numbe s o b i le s a s, so co als and glass sponges a e equalized esul ing in a mean a e age p ecision inc ease o 24 %. The o e all mean a e age p ecision o box de ec ions yields 76.7 and o ins ance segmen a ion 67.7 a an in e sec ion o e union h eshold o 0.5. This s udy shows ha deep a chi ec u es such as he deployed Cen e Mask ia ResNeX -101 model can success ully be ained wi h ew o iginal images om a ying unde wa e scenes. Keywo ds: objec de ec ion, deep lea ning, da a augmen a ion, ma ine image y, ben hic mega auna 1. INTRODUCTION Global al e a ion o ma ine en i onmen s due o o e ishing, pollu ion, physical habi a des uc ion and clima e change ha e led o an inc easing decline o animal di e si y and abundance h oughou many ma ine ecosys ems (Jackson, 2008). Especially ben hic mega auna in he Sou he n Ocean a e a isk o en i onmen al change and a e o signi ican ecological alue as hey al e small-scale opog aphy o seabed habi a s a ec ing he en i e ben hic communi y (Gili e al., 2006). Assessing he biodi e si y and cha ac e isa ion o ben hic communi ies is inc easingly impo an o iden i ying ulne able ma ine ecosys ems and de eloping conse a ion s a egies. Ma ine habi a s ha e been s udied based on mainly h ee p ac ices: physical sampling me hods using sledges o awling, acous ical echniques and op ical sys ems. While physical me hods a e able o collec specimen on a lowe axonomical scale hey des uc he en i onmen and hei sampling a e is a he low. Op ical sys ems a e mo e cos e ec i e, obus and p ecise han acous ic sys ems which led o a la ge g owing lib a y o digi al unde wa e images h oughou ecen yea s pa ing he way o new esea ch in au oma ic analy ical me hods. Objec de ec ion, classi ica ion and segmen a ion o image y has become a subs an ial ask in he ield o compu e ision. P io adi ional ea u e desc ip o s a e o en used o ex ac colou , shape o ex u e in o ma ion and a e good a de ec ing speci ic single objec s such as scallops (Dawkins e al., 2013) o lobs e s using classi ie s such as suppo ec o machine (Tan e al., 2018). Howe e , hey a e no obus o a ying unde wa e scenes ha a e exposed o ma ine snow, wa e u bidi y, lens dis o ion, spa se, uns able illumina ion and colou shi due o he su ey pla o m’s a ia ion in speed, angle o al i ude. Mo eo e , he e is a high a iabili y and in a iabili y be ween ea u es belonging o he same class and di e en classes, espec i ely (Pa oni e al., 2021). Deep lea ning (DL) using con olu ional neu al ne wo ks ha e p o en o ou pe o m adi ional based objec de ec ion (Gonzalez-Cid e al., 2017) as hey a e mo e in a ian o he de o ma ion o images. Addi ionally, images can di ec ly be used as inpu wi hou he necessi y o p e-p ocessing. Be e esul s a e o en achie ed using deepe laye s hus inc easing he numbe o pa ame e s o se e al millions. To achie e high pe o mances wi h hose models using images o uncons ained unde wa e scenes o ac oss a ying pla o ms, a la ge aining da ase size is he mos c ucial pa (Langenkämpe e al., 2020), o en, howe e , e y ime- consuming and cos ly o es ablish. This pape in es iga es he e ec o inpu image da a augmen a ion and composi ion s a egies in an a emp o o e come limi a ions o la ge da a se size and class imbalance. Taking accoun o he esul s, he s a e-o - he a ancho - ee objec de ec ion and ins ance segmen a ion model Cen e Mask (Lee and Pa k, 2019) ia ResNeX -101 (Xie e al., 2017) will be ained based on a small, highly di e se 700 image da ase . In his wo k, he de ec ion, segmen a ion and coun ing o glass sponges (hexac inellids), so co als (p imnoids and ch ysogo giids) and b i le s a s (ophiu oids) will be assessed p o iding i s s eps owa ds u u e abundance and size es ima ions o hese specimen. Deep Lea ning based De ec ion, Segmen a ion and Coun ing o Ben hic Mega auna in Uncons ained Unde wa e En i onmen s Mona Lü jens* Ha ald S e nbe g* *Ha enCi y Uni e si y Hambu g, Depa men o Hyd og aphy and Geodesy, Henning-Vosche au-Pla z 1, 20457 Hambu g, Ge many (e-mail: {mona.lue jens,ha ald.s e nbe g}@hcu-hambu g.de) Abs ac : Assessing and moni o ing ben hic communi ies is inc easingly impo an in iew o global al e a ion o ma ine en i onmen s. Deep lea ning has p o en o e ec i ely de ec ma ine specimen in unde wa e image y bu s ill ace p oblems wi h small inpu da ase s, uncons ained en i onmen s and class imbalance. This s udy e alua es a da a augmen a ion s a egy o alle ia e hese limi a ions. Th ough syn he ically de i ed image composi ions, he en i e inpu da ase was g ea ly ex ended om 700 o 12700 images. Addi ionally, specimen numbe s o b i le s a s, so co als and glass sponges a e equalized esul ing in a mean a e age p ecision inc ease o 24 %. The o e all mean a e age p ecision o box de ec ions yields 76.7 and o ins ance segmen a ion 67.7 a an in e sec ion o e union h eshold o 0.5. This s udy shows ha deep a chi ec u es such as he deployed Cen e Mask ia ResNeX -101 model can success ully be ained wi h ew o iginal images om a ying unde wa e scenes. Keywo ds: objec de ec ion, deep lea ning, da a augmen a ion, ma ine image y, ben hic mega auna 1. INTRODUCTION Global al e a ion o ma ine en i onmen s due o o e ishing, pollu ion, physical habi a des uc ion and clima e change ha e led o an inc easing decline o animal di e si y and abundance h oughou many ma ine ecosys ems (Jackson, 2008). Especially ben hic mega auna in he Sou he n Ocean a e a isk o en i onmen al change and a e o signi ican ecological alue as hey al e small-scale opog aphy o seabed habi a s a ec ing he en i e ben hic communi y (Gili e al., 2006). Assessing he biodi e si y and cha ac e isa ion o ben hic communi ies is inc easingly impo an o iden i ying ulne able ma ine ecosys ems and de eloping conse a ion s a egies. Ma ine habi a s ha e been s udied based on mainly h ee p ac ices: physical sampling me hods using sledges o awling, acous ical echniques and op ical sys ems. While physical me hods a e able o collec specimen on a lowe axonomical scale hey des uc he en i onmen and hei sampling a e is a he low. Op ical sys ems a e mo e cos e ec i e, obus and p ecise han acous ic sys ems which led o a la ge g owing lib a y o digi al unde wa e images h oughou ecen yea s pa ing he way o new esea ch in au oma ic analy ical me hods. Objec de ec ion, classi ica ion and segmen a ion o image y has become a subs an ial ask in he ield o compu e ision. P io adi ional ea u e desc ip o s a e o en used o ex ac colou , shape o ex u e in o ma ion and a e good a de ec ing speci ic single objec s such as scallops (Dawkins e al., 2013) o lobs e s using classi ie s such as suppo ec o machine (Tan e al., 2018). Howe e , hey a e no obus o a ying unde wa e scenes ha a e exposed o ma ine snow, wa e u bidi y, lens dis o ion, spa se, uns able illumina ion and colou shi due o he su ey pla o m’s a ia ion in speed, angle o al i ude. Mo eo e , he e is a high a iabili y and in a iabili y be ween ea u es belonging o he same class and di e en classes, espec i ely (Pa oni e al., 2021). Deep lea ning (DL) using con olu ional neu al ne wo ks ha e p o en o ou pe o m adi ional based objec de ec ion (Gonzalez-Cid e al., 2017) as hey a e mo e in a ian o he de o ma ion o images. Addi ionally, images can di ec ly be used as inpu wi hou he necessi y o p e-p ocessing. Be e esul s a e o en achie ed using deepe laye s hus inc easing he numbe o pa ame e s o se e al millions. To achie e high pe o mances wi h hose models using images o uncons ained unde wa e scenes o ac oss a ying pla o ms, a la ge aining da ase size is he mos c ucial pa (Langenkämpe e al., 2020), o en, howe e , e y ime- consuming and cos ly o es ablish. This pape in es iga es he e ec o inpu image da a augmen a ion and composi ion s a egies in an a emp o o e come limi a ions o la ge da a se size and class imbalance. Taking accoun o he esul s, he s a e-o - he a ancho - ee objec de ec ion and ins ance segmen a ion model Cen e Mask (Lee and Pa k, 2019) ia ResNeX -101 (Xie e al., 2017) will be ained based on a small, highly di e se 700 image da ase . In his wo k, he de ec ion, segmen a ion and coun ing o glass sponges (hexac inellids), so co als (p imnoids and ch ysogo giids) and b i le s a s (ophiu oids) will be assessed p o iding i s s eps owa ds u u e abundance and size es ima ions o hese specimen. Deep Lea ning based De ec ion, Segmen a ion and Coun ing o Ben hic Mega auna in Uncons ained Unde wa e En i onmen s Mona Lü jens* Ha ald S e nbe g* *Ha enCi y Uni e si y Hambu g, Depa men o Hyd og aphy and Geodesy, Henning-Vosche au-Pla z 1, 20457 Hambu g, Ge many (e-mail: {mona.lue jens,ha ald.s e nbe g}@hcu-hambu g.de) Abs ac : Assessing and moni o ing ben hic communi ies is inc easingly impo an in iew o global al e a ion o ma ine en i onmen s. Deep lea ning has p o en o e ec i ely de ec ma ine specimen in unde wa e image y bu s ill ace p oblems wi h small inpu da ase s, uncons ained en i onmen s and class imbalance. This s udy e alua es a da a augmen a ion s a egy o alle ia e hese limi a ions. Th ough syn he ically de i ed image composi ions, he en i e inpu da ase was g ea ly ex ended om 700 o 12700 images. Addi ionally, specimen numbe s o b i le s a s, so co als and glass sponges a e equalized esul ing in a mean a e age p ecision inc ease o 24 %. The o e all mean a e age p ecision o box de ec ions yields 76.7 and o ins ance segmen a ion 67.7 a an in e sec ion o e union h eshold o 0.5. This s udy shows ha deep a chi ec u es such as he deployed Cen e Mask ia ResNeX -101 model can success ully be ained wi h ew o iginal images om a ying unde wa e scenes. Keywo ds: objec de ec ion, deep lea ning, da a augmen a ion, ma ine image y, ben hic mega auna 1. INTRODUCTION Global al e a ion o ma ine en i onmen s due o o e ishing, pollu ion, physical habi a des uc ion and clima e change ha e led o an inc easing decline o animal di e si y and abundance h oughou many ma ine ecosys ems (Jackson, 2008). Especially ben hic mega auna in he Sou he n Ocean a e a isk o en i onmen al change and a e o signi ican ecological alue as hey al e small-scale opog aphy o seabed habi a s a ec ing he en i e ben hic communi y (Gili e al., 2006). Assessing he biodi e si y and cha ac e isa ion o ben hic communi ies is inc easingly impo an o iden i ying ulne able ma ine ecosys ems and de eloping conse a ion s a egies. Ma ine habi a s ha e been s udied based on mainly h ee p ac ices: physical sampling me hods using sledges o awling, acous ical echniques and op ical sys ems. While physical me hods a e able o collec specimen on a lowe axonomical scale hey des uc he en i onmen and hei sampling a e is a he low. Op ical sys ems a e mo e cos e ec i e, obus and p ecise han acous ic sys ems which led o a la ge g owing lib a y o digi al unde wa e images h oughou ecen yea s pa ing he way o new esea ch in au oma ic analy ical me hods. Objec de ec ion, classi ica ion and segmen a ion o image y has become a subs an ial ask in he ield o compu e ision. P io adi ional ea u e desc ip o s a e o en used o ex ac colou , shape o ex u e in o ma ion and a e good a de ec ing speci ic single objec s such as scallops (Dawkins e al., 2013) o lobs e s using classi ie s such as suppo ec o machine (Tan e al., 2018). Howe e , hey a e no obus o a ying unde wa e scenes ha a e exposed o ma ine snow, wa e u bidi y, lens dis o ion, spa se, uns able illumina ion and colou shi due o he su ey pla o m’s a ia ion in speed, angle o al i ude. Mo eo e , he e is a high a iabili y and in a iabili y be ween ea u es belonging o he same class and di e en classes, espec i ely (Pa oni e al., 2021). Deep lea ning (DL) using con olu ional neu al ne wo ks ha e p o en o ou pe o m adi ional based objec de ec ion (Gonzalez-Cid e al., 2017) as hey a e mo e in a ian o he de o ma ion o images. Addi ionally, images can di ec ly be used as inpu wi hou he necessi y o p e-p ocessing. Be e esul s a e o en achie ed using deepe laye s hus inc easing he numbe o pa ame e s o se e al millions. To achie e high pe o mances wi h hose models using images o uncons ained unde wa e scenes o ac oss a ying pla o ms, a la ge aining da ase size is he mos c ucial pa (Langenkämpe e al., 2020), o en, howe e , e y ime- consuming and cos ly o es ablish. This pape in es iga es he e ec o inpu image da a augmen a ion and composi ion s a egies in an a emp o o e come limi a ions o la ge da a se size and class imbalance. Taking accoun o he esul s, he s a e-o - he a ancho - ee objec de ec ion and ins ance segmen a ion model Cen e Mask (Lee and Pa k, 2019) ia ResNeX -101 (Xie e al., 2017) will be ained based on a small, highly di e se 700 image da ase . In his wo k, he de ec ion, segmen a ion and coun ing o glass sponges (hexac inellids), so co als (p imnoids and ch ysogo giids) and b i le s a s (ophiu oids) will be assessed p o iding i s s eps owa ds u u e abundance and size es ima ions o hese specimen. Deep Lea ning based De ec ion, Segmen a ion and Coun ing o Ben hic Mega auna in Uncons ained Unde wa e En i onmen s Mona Lü jens* Ha ald S e nbe g* *Ha enCi y Uni e si y Hambu g, Depa men o Hyd og aphy and Geodesy, Henning-Vosche au-Pla z 1, 20457 Hambu g, Ge many (e-mail: {mona.lue jens,ha ald.s e nbe g}@hcu-hambu g.de) Abs ac : Assessing and moni o ing ben hic communi ies is inc easingly impo an in iew o global al e a ion o ma ine en i onmen s. Deep lea ning has p o en o e ec i ely de ec ma ine specimen in unde wa e image y bu s ill ace p oblems wi h small inpu da ase s, uncons ained en i onmen s and class imbalance. This s udy e alua es a da a augmen a ion s a egy o alle ia e hese limi a ions. Th ough syn he ically de i ed image composi ions, he en i e inpu da ase was g ea ly ex ended om 700 o 12700 images. Addi ionally, specimen numbe s o b i le s a s, so co als and glass sponges a e equalized esul ing in a mean a e age p ecision inc ease o 24 %. The o e all mean a e age p ecision o box de ec ions yields 76.7 and o ins ance segmen a ion 67.7 a an in e sec ion o e union h eshold o 0.5. This s udy shows ha deep a chi ec u es such as he deployed Cen e Mask ia ResNeX -101 model can success ully be ained wi h ew o iginal images om a ying unde wa e scenes. Keywo ds: objec de ec ion, deep lea ning, da a augmen a ion, ma ine image y, ben hic mega auna 1. INTRODUCTION Global al e a ion o ma ine en i onmen s due o o e ishing, pollu ion, physical habi a des uc ion and clima e change ha e led o an inc easing decline o animal di e si y and abundance h oughou many ma ine ecosys ems (Jackson, 2008). Especially ben hic mega auna in he Sou he n Ocean a e a isk o en i onmen al change and a e o signi ican ecological alue as hey al e small-scale opog aphy o seabed habi a s a ec ing he en i e ben hic communi y (Gili e al., 2006). Assessing he biodi e si y and cha ac e isa ion o ben hic communi ies is inc easingly impo an o iden i ying ulne able ma ine ecosys ems and de eloping conse a ion s a egies. Ma ine habi a s ha e been s udied based on mainly h ee p ac ices: physical sampling me hods using sledges o awling, acous ical echniques and op ical sys ems. While physical me hods a e able o collec specimen on a lowe axonomical scale hey des uc he en i onmen and hei sampling a e is a he low. Op ical sys ems a e mo e cos e ec i e, obus and p ecise han acous ic sys ems which led o a la ge g owing lib a y o digi al unde wa e images h oughou ecen yea s pa ing he way o new esea ch in au oma ic analy ical me hods. Objec de ec ion, classi ica ion and segmen a ion o image y has become a subs an ial ask in he ield o compu e ision. P io adi ional ea u e desc ip o s a e o en used o ex ac colou , shape o ex u e in o ma ion and a e good a de ec ing speci ic single objec s such as scallops (Dawkins e al., 2013) o lobs e s using classi ie s such as suppo ec o machine (Tan e al., 2018). Howe e , hey a e no obus o a ying unde wa e scenes ha a e exposed o ma ine snow, wa e u bidi y, lens dis o ion, spa se, uns able illumina ion and colou shi due o he su ey pla o m’s a ia ion in speed, angle o al i ude. Mo eo e , he e is a high a iabili y and in a iabili y be ween ea u es belonging o he same class and di e en classes, espec i ely (Pa oni e al., 2021). Deep lea ning (DL) using con olu ional neu al ne wo ks ha e p o en o ou pe o m adi ional based objec de ec ion (Gonzalez-Cid e al., 2017) as hey a e mo e in a ian o he de o ma ion o images. Addi ionally, images can di ec ly be used as inpu wi hou he necessi y o p e-p ocessing. Be e esul s a e o en achie ed using deepe laye s hus inc easing he numbe o pa ame e s o se e al millions. To achie e high pe o mances wi h hose models using images o uncons ained unde wa e scenes o ac oss a ying pla o ms, a la ge aining da ase size is he mos c ucial pa (Langenkämpe e al., 2020), o en, howe e , e y ime- consuming and cos ly o es ablish. This pape in es iga es he e ec o inpu image da a augmen a ion and composi ion s a egies in an a emp o o e come limi a ions o la ge da a se size and class imbalance. Taking accoun o he esul s, he s a e-o - he a ancho - ee objec de ec ion and ins ance segmen a ion model Cen e Mask (Lee and Pa k, 2019) ia ResNeX -101 (Xie e al., 2017) will be ained based on a small, highly di e se 700 image da ase . In his wo k, he de ec ion, segmen a ion and coun ing o glass sponges (hexac inellids), so co als (p imnoids and ch ysogo giids) and b i le s a s (ophiu oids) will be assessed p o iding i s s eps owa ds u u e abundance and size es ima ions o hese specimen. Deep Lea ning based De ec ion, Segmen a ion and Coun ing o Ben hic Mega auna in Uncons ained Unde wa e En i onmen s Mona Lü jens* Ha ald S e nbe g* *Ha enCi y Uni e si y Hambu g, Depa men o Hyd og aphy and Geodesy, Henning-Vosche au-Pla z 1, 20457 Hambu g, Ge many (e-mail: {mona.lue jens,ha ald.s e nbe g}@hcu-hambu g.de) Abs ac : Assessing and moni o ing ben hic communi ies is inc easingly impo an in iew o global al e a ion o ma ine en i onmen s. Deep lea ning has p o en o e ec i ely de ec ma ine specimen in unde wa e image y bu s ill ace p oblems wi h small inpu da ase s, uncons ained en i onmen s and class imbalance. This s udy e alua es a da a augmen a ion s a egy o alle ia e hese limi a ions. Th ough syn he ically de i ed image composi ions, he en i e inpu da ase was g ea ly ex ended om 700 o 12700 images. Addi ionally, specimen numbe s o b i le s a s, so co als and glass sponges a e equalized esul ing in a mean a e age p ecision inc ease o 24 %. The o e all mean a e age p ecision o box de ec ions yields 76.7 and o ins ance segmen a ion 67.7 a an in e sec ion o e union h eshold o 0.5. This s udy shows ha deep a chi ec u es such as he deployed Cen e Mask ia ResNeX -101 model can success ully be ained wi h ew o iginal images om a ying unde wa e scenes. Keywo ds: objec de ec ion, deep lea ning, da a augmen a ion, ma ine image y, ben hic mega auna 1. INTRODUCTION Global al e a ion o ma ine en i onmen s due o o e ishing, pollu ion, physical habi a des uc ion and clima e change ha e led o an inc easing decline o animal di e si y and abundance h oughou many ma ine ecosys ems (Jackson, 2008). Especially ben hic mega auna in he Sou he n Ocean a e a isk o en i onmen al change and a e o signi ican ecological alue as hey al e small-scale opog aphy o seabed habi a s a ec ing he en i e ben hic communi y (Gili e al., 2006). Assessing he biodi e si y and cha ac e isa ion o ben hic communi ies is inc easingly impo an o iden i ying ulne able ma ine ecosys ems and de eloping conse a ion s a egies. Ma ine habi a s ha e been s udied based on mainly h ee p ac ices: physical sampling me hods using sledges o awling, acous ical echniques and op ical sys ems. While physical me hods a e able o collec specimen on a lowe axonomical scale hey des uc he en i onmen and hei sampling a e is a he low. Op ical sys ems a e mo e cos e ec i e, obus and p ecise han acous ic sys ems which led o a la ge g owing lib a y o digi al unde wa e images h oughou ecen yea s pa ing he way o new esea ch in au oma ic analy ical me hods. Objec de ec ion, classi ica ion and segmen a ion o image y has become a subs an ial ask in he ield o compu e ision. P io adi ional ea u e desc ip o s a e o en used o ex ac colou , shape o ex u e in o ma ion and a e good a de ec ing speci ic single objec s such as scallops (Dawkins e al., 2013) o lobs e s using classi ie s such as suppo ec o machine (Tan e al., 2018). Howe e , hey a e no obus o a ying unde wa e scenes ha a e exposed o ma ine snow, wa e u bidi y, lens dis o ion, spa se, uns able illumina ion and colou shi due o he su ey pla o m’s a ia ion in speed, angle o al i ude. Mo eo e , he e is a high a iabili y and in a iabili y be ween ea u es belonging o he same class and di e en classes, espec i ely (Pa oni e al., 2021). Deep lea ning (DL) using con olu ional neu al ne wo ks ha e p o en o ou pe o m adi ional based objec de ec ion (Gonzalez-Cid e al., 2017) as hey a e mo e in a ian o he de o ma ion o images. Addi ionally, images can di ec ly be used as inpu wi hou he necessi y o p e-p ocessing. Be e esul s a e o en achie ed using deepe laye s hus inc easing he numbe o pa ame e s o se e al millions. To achie e high pe o mances wi h hose models using images o uncons ained unde wa e scenes o ac oss a ying pla o ms, a la ge aining da ase size is he mos c ucial pa (Langenkämpe e al., 2020), o en, howe e , e y ime- consuming and cos ly o es ablish. This pape in es iga es he e ec o inpu image da a augmen a ion and composi ion s a egies in an a emp o o e come limi a ions o la ge da a se size and class imbalance. Taking accoun o he esul s, he s a e-o - he a ancho - ee objec de ec ion and ins ance segmen a ion model Cen e Mask (Lee and Pa k, 2019) ia ResNeX -101 (Xie e al., 2017) will be ained based on a small, highly di e se 700 image da ase . In his wo k, he de ec ion, segmen a ion and coun ing o glass sponges (hexac inellids), so co als (p imnoids and ch ysogo giids) and b i le s a s (ophiu oids) will be assessed p o iding i s s eps owa ds u u e abundance and size es ima ions o hese specimen. Mona Lü jens e al. / IFAC Pape sOnLine 54-16 (2021) 76–82 77 Copy igh © 2021 The Au ho s. This is an open access a icle unde he CC BY-NC-ND license ( h ps://c ea i ecommons.o g/licenses/by-nc-nd/4.0/ ) Deep Lea ning based De ec ion, Segmen a ion and Coun ing o Ben hic Mega auna in Uncons ained Unde wa e En i onmen s Mona Lü jens* Ha ald S e nbe g* *Ha enCi y Uni e si y Hambu g, Depa men o Hyd og aphy and Geodesy, Henning-Vosche au-Pla z 1, 20457 Hambu g, Ge many (e-mail: {mona.lue jens,ha ald.s e nbe g}@hcu-hambu g.de) Abs ac : Assessing and moni o ing ben hic communi ies is inc easingly impo an in iew o global al e a ion o ma ine en i onmen s. Deep lea ning has p o en o e ec i ely de ec ma ine specimen in unde wa e image y bu s ill ace p oblems wi h small inpu da ase s, uncons ained en i onmen s and class imbalance. This s udy e alua es a da a augmen a ion s a egy o alle ia e hese limi a ions. Th ough syn he ically de i ed image composi ions, he en i e inpu da ase was g ea ly ex ended om 700 o 12700 images. Addi ionally, specimen numbe s o b i le s a s, so co als and glass sponges a e equalized esul ing in a mean a e age p ecision inc ease o 24 %. The o e all mean a e age p ecision o box de ec ions yields 76.7 and o ins ance segmen a ion 67.7 a an in e sec ion o e union h eshold o 0.5. This s udy shows ha deep a chi ec u es such as he deployed Cen e Mask ia ResNeX -101 model can success ully be ained wi h ew o iginal images om a ying unde wa e scenes. Keywo ds: objec de ec ion, deep lea ning, da a augmen a ion, ma ine image y, ben hic mega auna 1. INTRODUCTION Global al e a ion o ma ine en i onmen s due o o e ishing, pollu ion, physical habi a des uc ion and clima e change ha e led o an inc easing decline o animal di e si y and abundance h oughou many ma ine ecosys ems (Jackson, 2008). Especially ben hic mega auna in he Sou he n Ocean a e a isk o en i onmen al change and a e o signi ican ecological alue as hey al e small-scale opog aphy o seabed habi a s a ec ing he en i e ben hic communi y (Gili e al., 2006). Assessing he biodi e si y and cha ac e isa ion o ben hic communi ies is inc easingly impo an o iden i ying ulne able ma ine ecosys ems and de eloping conse a ion s a egies. Ma ine habi a s ha e been s udied based on mainly h ee p ac ices: physical sampling me hods using sledges o awling, acous ical echniques and op ical sys ems. While physical me hods a e able o collec specimen on a lowe axonomical scale hey des uc he en i onmen and hei sampling a e is a he low. Op ical sys ems a e mo e cos e ec i e, obus and p ecise han acous ic sys ems which led o a la ge g owing lib a y o digi al unde wa e images h oughou ecen yea s pa ing he way o new esea ch in au oma ic analy ical me hods. Objec de ec ion, classi ica ion and segmen a ion o image y has become a subs an ial ask in he ield o compu e ision. P io adi ional ea u e desc ip o s a e o en used o ex ac colou , shape o ex u e in o ma ion and a e good a de ec ing speci ic single objec s such as scallops (Dawkins e al., 2013) o lobs e s using classi ie s such as suppo ec o machine (Tan e al., 2018). Howe e , hey a e no obus o a ying unde wa e scenes ha a e exposed o ma ine snow, wa e u bidi y, lens dis o ion, spa se, uns able illumina ion and colou shi due o he su ey pla o m’s a ia ion in speed, angle o al i ude. Mo eo e , he e is a high a iabili y and in a iabili y be ween ea u es belonging o he same class and di e en classes, espec i ely (Pa oni e al., 2021). Deep lea ning (DL) using con olu ional neu al ne wo ks ha e p o en o ou pe o m adi ional based objec de ec ion (Gonzalez-Cid e al., 2017) as hey a e mo e in a ian o he de o ma ion o images. Addi ionally, images can di ec ly be used as inpu wi hou he necessi y o p e-p ocessing. Be e esul s a e o en achie ed using deepe laye s hus inc easing he numbe o pa ame e s o se e al millions. To achie e high pe o mances wi h hose models using images o uncons ained unde wa e scenes o ac oss a ying pla o ms, a la ge aining da ase size is he mos c ucial pa (Langenkämpe e al., 2020), o en, howe e , e y ime- consuming and cos ly o es ablish. This pape in es iga es he e ec o inpu image da a augmen a ion and composi ion s a egies in an a emp o o e come limi a ions o la ge da a se size and class imbalance. Taking accoun o he esul s, he s a e-o - he a ancho - ee objec de ec ion and ins ance segmen a ion model Cen e Mask (Lee and Pa k, 2019) ia ResNeX -101 (Xie e al., 2017) will be ained based on a small, highly di e se 700 image da ase . In his wo k, he de ec ion, segmen a ion and coun ing o glass sponges (hexac inellids), so co als (p imnoids and ch ysogo giids) and b i le s a s (ophiu oids) will be assessed p o iding i s s eps owa ds u u e abundance and size es ima ions o hese specimen. Deep Lea ning based De ec ion, Segmen a ion and Coun ing o Ben hic Mega auna in Uncons ained Unde wa e En i onmen s Mona Lü jens* Ha ald S e nbe g* *Ha enCi y Uni e si y Hambu g, Depa men o Hyd og aphy and Geodesy, Henning-Vosche au-Pla z 1, 20457 Hambu g, Ge many (e-mail: {mona.lue jens,ha ald.s e nbe g}@hcu-hambu g.de) Abs ac : Assessing and moni o ing ben hic communi ies is inc easingly impo an in iew o global al e a ion o ma ine en i onmen s. Deep lea ning has p o en o e ec i ely de ec ma ine specimen in unde wa e image y bu s ill ace p oblems wi h small inpu da ase s, uncons ained en i onmen s and class imbalance. This s udy e alua es a da a augmen a ion s a egy o alle ia e hese limi a ions. Th ough syn he ically de i ed image composi ions, he en i e inpu da ase was g ea ly ex ended om 700 o 12700 images. Addi ionally, specimen numbe s o b i le s a s, so co als and glass sponges a e equalized esul ing in a mean a e age p ecision inc ease o 24 %. The o e all mean a e age p ecision o box de ec ions yields 76.7 and o ins ance segmen a ion 67.7 a an in e sec ion o e union h eshold o 0.5. This s udy shows ha deep a chi ec u es such as he deployed Cen e Mask ia ResNeX -101 model can success ully be ained wi h ew o iginal images om a ying unde wa e scenes. Keywo ds: objec de ec ion, deep lea ning, da a augmen a ion, ma ine image y, ben hic mega auna 1. INTRODUCTION Global al e a ion o ma ine en i onmen s due o o e ishing, pollu ion, physical habi a des uc ion and clima e change ha e led o an inc easing decline o animal di e si y and abundance h oughou many ma ine ecosys ems (Jackson, 2008). Especially ben hic mega auna in he Sou he n Ocean a e a isk o en i onmen al change and a e o signi ican ecological alue as hey al e small-scale opog aphy o seabed habi a s a ec ing he en i e ben hic communi y (Gili e al., 2006). Assessing he biodi e si y and cha ac e isa ion o ben hic communi ies is inc easingly impo an o iden i ying ulne able ma ine ecosys ems and de eloping conse a ion s a egies. Ma ine habi a s ha e been s udied based on mainly h ee p ac ices: physical sampling me hods using sledges o awling, acous ical echniques and op ical sys ems. While physical me hods a e able o collec specimen on a lowe axonomical scale hey des uc he en i onmen and hei sampling a e is a he low. Op ical sys ems a e mo e cos e ec i e, obus and p ecise han acous ic sys ems which led o a la ge g owing lib a y o digi al unde wa e images h oughou ecen yea s pa ing he way o new esea ch in au oma ic analy ical me hods. Objec de ec ion, classi ica ion and segmen a ion o image y has become a subs an ial ask in he ield o compu e ision. P io adi ional ea u e desc ip o s a e o en used o ex ac colou , shape o ex u e in o ma ion and a e good a de ec ing speci ic single objec s such as scallops (Dawkins e al., 2013) o lobs e s using classi ie s such as suppo ec o machine (Tan e al., 2018). Howe e , hey a e no obus o a ying unde wa e scenes ha a e exposed o ma ine snow, wa e u bidi y, lens dis o ion, spa se, uns able illumina ion and colou shi due o he su ey pla o m’s a ia ion in speed, angle o al i ude. Mo eo e , he e is a high a iabili y and in a iabili y be ween ea u es belonging o he same class and di e en classes, espec i ely (Pa oni e al., 2021). Deep lea ning (DL) using con olu ional neu al ne wo ks ha e p o en o ou pe o m adi ional based objec de ec ion (Gonzalez-Cid e al., 2017) as hey a e mo e in a ian o he de o ma ion o images. Addi ionally, images can di ec ly be used as inpu wi hou he necessi y o p e-p ocessing. Be e esul s a e o en achie ed using deepe laye s hus inc easing he numbe o pa ame e s o se e al millions. To achie e high pe o mances wi h hose models using images o uncons ained unde wa e scenes o ac oss a ying pla o ms, a la ge aining da ase size is he mos c ucial pa (Langenkämpe e al., 2020), o en, howe e , e y ime- consuming and cos ly o es ablish. This pape in es iga es he e ec o inpu image da a augmen a ion and composi ion s a egies in an a emp o o e come limi a ions o la ge da a se size and class imbalance. Taking accoun o he esul s, he s a e-o - he a ancho - ee objec de ec ion and ins ance segmen a ion model Cen e Mask (Lee and Pa k, 2019) ia ResNeX -101 (Xie e al., 2017) will be ained based on a small, highly di e se 700 image da ase . In his wo k, he de ec ion, segmen a ion and coun ing o glass sponges (hexac inellids), so co als (p imnoids and ch ysogo giids) and b i le s a s (ophiu oids) will be assessed p o iding i s s eps owa ds u u e abundance and size es ima ions o hese specimen. Deep Lea ning based De ec ion, Segmen a ion and Coun ing o Ben hic Mega auna in Uncons ained Unde wa e En i onmen s Mona Lü jens* Ha ald S e nbe g* *Ha enCi y Uni e si y Hambu g, Depa men o Hyd og aphy and Geodesy, Henning-Vosche au-Pla z 1, 20457 Hambu g, Ge many (e-mail: {mona.lue jens,ha ald.s e nbe g}@hcu-hambu g.de) Abs ac : Assessing and moni o ing ben hic communi ies is inc easingly impo an in iew o global al e a ion o ma ine en i onmen s. Deep lea ning has p o en o e ec i ely de ec ma ine specimen in unde wa e image y bu s ill ace p oblems wi h small inpu da ase s, uncons ained en i onmen s and class imbalance. This s udy e alua es a da a augmen a ion s a egy o alle ia e hese limi a ions. Th ough syn he ically de i ed image composi ions, he en i e inpu da ase was g ea ly ex ended om 700 o 12700 images. Addi ionally, specimen numbe s o b i le s a s, so co als and glass sponges a e equalized esul ing in a mean a e age p ecision inc ease o 24 %. The o e all mean a e age p ecision o box de ec ions yields 76.7 and o ins ance segmen a ion 67.7 a an in e sec ion o e union h eshold o 0.5. This s udy shows ha deep a chi ec u es such as he deployed Cen e Mask ia ResNeX -101 model can success ully be ained wi h ew o iginal images om a ying unde wa e scenes. Keywo ds: objec de ec ion, deep lea ning, da a augmen a ion, ma ine image y, ben hic mega auna 1. INTRODUCTION Global al e a ion o ma ine en i onmen s due o o e ishing, pollu ion, physical habi a des uc ion and clima e change ha e led o an inc easing decline o animal di e si y and abundance h oughou many ma ine ecosys ems (Jackson, 2008). Especially ben hic mega auna in he Sou he n Ocean a e a isk o en i onmen al change and a e o signi ican ecological alue as hey al e small-scale opog aphy o seabed habi a s a ec ing he en i e ben hic communi y (Gili e al., 2006). Assessing he biodi e si y and cha ac e isa ion o ben hic communi ies is inc easingly impo an o iden i ying ulne able ma ine ecosys ems and de eloping conse a ion s a egies. Ma ine habi a s ha e been s udied based on mainly h ee p ac ices: physical sampling me hods using sledges o awling, acous ical echniques and op ical sys ems. While physical me hods a e able o collec specimen on a lowe axonomical scale hey des uc he en i onmen and hei sampling a e is a he low. Op ical sys ems a e mo e cos e ec i e, obus and p ecise han acous ic sys ems which led o a la ge g owing lib a y o digi al unde wa e images h oughou ecen yea s pa ing he way o new esea ch in au oma ic analy ical me hods. Objec de ec ion, classi ica ion and segmen a ion o image y has become a subs an ial ask in he ield o compu e ision. P io adi ional ea u e desc ip o s a e o en used o ex ac colou , shape o ex u e in o ma ion and a e good a de ec ing speci ic single objec s such as scallops (Dawkins e al., 2013) o lobs e s using classi ie s such as suppo ec o machine (Tan e al., 2018). Howe e , hey a e no obus o a ying unde wa e scenes ha a e exposed o ma ine snow, wa e u bidi y, lens dis o ion, spa se, uns able illumina ion and colou shi due o he su ey pla o m’s a ia ion in speed, angle o al i ude. Mo eo e , he e is a high a iabili y and in a iabili y be ween ea u es belonging o he same class and di e en classes, espec i ely (Pa oni e al., 2021). Deep lea ning (DL) using con olu ional neu al ne wo ks ha e p o en o ou pe o m adi ional based objec de ec ion (Gonzalez-Cid e al., 2017) as hey a e mo e in a ian o he de o ma ion o images. Addi ionally, images can di ec ly be used as inpu wi hou he necessi y o p e-p ocessing. Be e esul s a e o en achie ed using deepe laye s hus inc easing he numbe o pa ame e s o se e al millions. To achie e high pe o mances wi h hose models using images o uncons ained unde wa e scenes o ac oss a ying pla o ms, a la ge aining da ase size is he mos c ucial pa (Langenkämpe e al., 2020), o en, howe e , e y ime- consuming and cos ly o es ablish. This pape in es iga es he e ec o inpu image da a augmen a ion and composi ion s a egies in an a emp o o e come limi a ions o la ge da a se size and class imbalance. Taking accoun o he esul s, he s a e-o - he a ancho - ee objec de ec ion and ins ance segmen a ion model Cen e Mask (Lee and Pa k, 2019) ia ResNeX -101 (Xie e al., 2017) will be ained based on a small, highly di e se 700 image da ase . In his wo k, he de ec ion, segmen a ion and coun ing o glass sponges (hexac inellids), so co als (p imnoids and ch ysogo giids) and b i le s a s (ophiu oids) will be assessed p o iding i s s eps owa ds u u e abundance and size es ima ions o hese specimen. Deep Lea ning based De ec ion, Segmen a ion and Coun ing o Ben hic Mega auna in Uncons ained Unde wa e En i onmen s Mona Lü jens* Ha ald S e nbe g* *Ha enCi y Uni e si y Hambu g, Depa men o Hyd og aphy and Geodesy, Henning-Vosche au-Pla z 1, 20457 Hambu g, Ge many (e-mail: {mona.lue jens,ha ald.s e nbe g}@hcu-hambu g.de) Abs ac : Assessing and moni o ing ben hic communi ies is inc easingly impo an in iew o global al e a ion o ma ine en i onmen s. Deep lea ning has p o en o e ec i ely de ec ma ine specimen in unde wa e image y bu s ill ace p oblems wi h small inpu da ase s, uncons ained en i onmen s and class imbalance. This s udy e alua es a da a augmen a ion s a egy o alle ia e hese limi a ions. Th ough syn he ically de i ed image composi ions, he en i e inpu da ase was g ea ly ex ended om 700 o 12700 images. Addi ionally, specimen numbe s o b i le s a s, so co als and glass sponges a e equalized esul ing in a mean a e age p ecision inc ease o 24 %. The o e all mean a e age p ecision o box de ec ions yields 76.7 and o ins ance segmen a ion 67.7 a an in e sec ion o e union h eshold o 0.5. This s udy shows ha deep a chi ec u es such as he deployed Cen e Mask ia ResNeX -101 model can success ully be ained wi h ew o iginal images om a ying unde wa e scenes. Keywo ds: objec de ec ion, deep lea ning, da a augmen a ion, ma ine image y, ben hic mega auna 1. INTRODUCTION Global al e a ion o ma ine en i onmen s due o o e ishing, pollu ion, physical habi a des uc ion and clima e change ha e led o an inc easing decline o animal di e si y and abundance h oughou many ma ine ecosys ems (Jackson, 2008). Especially ben hic mega auna in he Sou he n Ocean a e a isk o en i onmen al change and a e o signi ican ecological alue as hey al e small-scale opog aphy o seabed habi a s a ec ing he en i e ben hic communi y (Gili e al., 2006). Assessing he biodi e si y and cha ac e isa ion o ben hic communi ies is inc easingly impo an o iden i ying ulne able ma ine ecosys ems and de eloping conse a ion s a egies. Ma ine habi a s ha e been s udied based on mainly h ee p ac ices: physical sampling me hods using sledges o awling, acous ical echniques and op ical sys ems. While physical me hods a e able o collec specimen on a lowe axonomical scale hey des uc he en i onmen and hei sampling a e is a he low. Op ical sys ems a e mo e cos e ec i e, obus and p ecise han acous ic sys ems which led o a la ge g owing lib a y o digi al unde wa e images h oughou ecen yea s pa ing he way o new esea ch in au oma ic analy ical me hods. Objec de ec ion, classi ica ion and segmen a ion o image y has become a subs an ial ask in he ield o compu e ision. P io adi ional ea u e desc ip o s a e o en used o ex ac colou , shape o ex u e in o ma ion and a e good a de ec ing speci ic single objec s such as scallops (Dawkins e al., 2013) o lobs e s using classi ie s such as suppo ec o machine (Tan e al., 2018). Howe e , hey a e no obus o a ying unde wa e scenes ha a e exposed o ma ine snow, wa e u bidi y, lens dis o ion, spa se, uns able illumina ion and colou shi due o he su ey pla o m’s a ia ion in speed, angle o al i ude. Mo eo e , he e is a high a iabili y and in a iabili y be ween ea u es belonging o he same class and di e en classes, espec i ely (Pa oni e al., 2021). Deep lea ning (DL) using con olu ional neu al ne wo ks ha e p o en o ou pe o m adi ional based objec de ec ion (Gonzalez-Cid e al., 2017) as hey a e mo e in a ian o he de o ma ion o images. Addi ionally, images can di ec ly be used as inpu wi hou he necessi y o p e-p ocessing. Be e esul s a e o en achie ed using deepe laye s hus inc easing he numbe o pa ame e s o se e al millions. To achie e high pe o mances wi h hose models using images o uncons ained unde wa e scenes o ac oss a ying pla o ms, a la ge aining da ase size is he mos c ucial pa (Langenkämpe e al., 2020), o en, howe e , e y ime- consuming and cos ly o es ablish. This pape in es iga es he e ec o inpu image da a augmen a ion and composi ion s a egies in an a emp o o e come limi a ions o la ge da a se size and class imbalance. Taking accoun o he esul s, he s a e-o - he a ancho - ee objec de ec ion and ins ance segmen a ion model Cen e Mask (Lee and Pa k, 2019) ia ResNeX -101 (Xie e al., 2017) will be ained based on a small, highly di e se 700 image da ase . In his wo k, he de ec ion, segmen a ion and coun ing o glass sponges (hexac inellids), so co als (p imnoids and ch ysogo giids) and b i le s a s (ophiu oids) will be assessed p o iding i s s eps owa ds u u e abundance and size es ima ions o hese specimen. Deep Lea ning based De ec ion, Segmen a ion and Coun ing o Ben hic Mega auna in Uncons ained Unde wa e En i onmen s Mona Lü jens* Ha ald S e nbe g* *Ha enCi y Uni e si y Hambu g, Depa men o Hyd og aphy and Geodesy, Henning-Vosche au-Pla z 1, 20457 Hambu g, Ge many (e-mail: {mona.lue jens,ha ald.s e nbe g}@hcu-hambu g.de) Abs ac : Assessing and moni o ing ben hic communi ies is inc easingly impo an in iew o global al e a ion o ma ine en i onmen s. Deep lea ning has p o en o e ec i ely de ec ma ine specimen in unde wa e image y bu s ill ace p oblems wi h small inpu da ase s, uncons ained en i onmen s and class imbalance. This s udy e alua es a da a augmen a ion s a egy o alle ia e hese limi a ions. Th ough syn he ically de i ed image composi ions, he en i e inpu da ase was g ea ly ex ended om 700 o 12700 images. Addi ionally, specimen numbe s o b i le s a s, so co als and glass sponges a e equalized esul ing in a mean a e age p ecision inc ease o 24 %. The o e all mean a e age p ecision o box de ec ions yields 76.7 and o ins ance segmen a ion 67.7 a an in e sec ion o e union h eshold o 0.5. This s udy shows ha deep a chi ec u es such as he deployed Cen e Mask ia ResNeX -101 model can success ully be ained wi h ew o iginal images om a ying unde wa e scenes. Keywo ds: objec de ec ion, deep lea ning, da a augmen a ion, ma ine image y, ben hic mega auna 1. INTRODUCTION Global al e a ion o ma ine en i onmen s due o o e ishing, pollu ion, physical habi a des uc ion and clima e change ha e led o an inc easing decline o animal di e si y and abundance h oughou many ma ine ecosys ems (Jackson, 2008). Especially ben hic mega auna in he Sou he n Ocean a e a isk o en i onmen al change and a e o signi ican ecological alue as hey al e small-scale opog aphy o seabed habi a s a ec ing he en i e ben hic communi y (Gili e al., 2006). Assessing he biodi e si y and cha ac e isa ion o ben hic communi ies is inc easingly impo an o iden i ying ulne able ma ine ecosys ems and de eloping conse a ion s a egies. Ma ine habi a s ha e been s udied based on mainly h ee p ac ices: physical sampling me hods using sledges o awling, acous ical echniques and op ical sys ems. While physical me hods a e able o collec specimen on a lowe axonomical scale hey des uc he en i onmen and hei sampling a e is a he low. Op ical sys ems a e mo e cos e ec i e, obus and p ecise han acous ic sys ems which led o a la ge g owing lib a y o digi al unde wa e images h oughou ecen yea s pa ing he way o new esea ch in au oma ic analy ical me hods. Objec de ec ion, classi ica ion and segmen a ion o image y has become a subs an ial ask in he ield o compu e ision. P io adi ional ea u e desc ip o s a e o en used o ex ac colou , shape o ex u e in o ma ion and a e good a de ec ing speci ic single objec s such as scallops (Dawkins e al., 2013) o lobs e s using classi ie s such as suppo ec o machine (Tan e al., 2018). Howe e , hey a e no obus o a ying unde wa e scenes ha a e exposed o ma ine snow, wa e u bidi y, lens dis o ion, spa se, uns able illumina ion and colou shi due o he su ey pla o m’s a ia ion in speed, angle o al i ude. Mo eo e , he e is a high a iabili y and in a iabili y be ween ea u es belonging o he same class and di e en classes, espec i ely (Pa oni e al., 2021). Deep lea ning (DL) using con olu ional neu al ne wo ks ha e p o en o ou pe o m adi ional based objec de ec ion (Gonzalez-Cid e al., 2017) as hey a e mo e in a ian o he de o ma ion o images. Addi ionally, images can di ec ly be used as inpu wi hou he necessi y o p e-p ocessing. Be e esul s a e o en achie ed using deepe laye s hus inc easing he numbe o pa ame e s o se e al millions. To achie e high pe o mances wi h hose models using images o uncons ained unde wa e scenes o ac oss a ying pla o ms, a la ge aining da ase size is he mos c ucial pa (Langenkämpe e al., 2020), o en, howe e , e y ime- consuming and cos ly o es ablish. This pape in es iga es he e ec o inpu image da a augmen a ion and composi ion s a egies in an a emp o o e come limi a ions o la ge da a se size and class imbalance. Taking accoun o he esul s, he s a e-o - he a ancho - ee objec de ec ion and ins ance segmen a ion model Cen e Mask (Lee and Pa k, 2019) ia ResNeX -101 (Xie e al., 2017) will be ained based on a small, highly di e se 700 image da ase . In his wo k, he de ec ion, segmen a ion and coun ing o glass sponges (hexac inellids), so co als (p imnoids and ch ysogo giids) and b i le s a s (ophiu oids) will be assessed p o iding i s s eps owa ds u u e abundance and size es ima ions o hese specimen. Deep Lea ning based De ec ion, Segmen a ion and Coun ing o Ben hic Mega auna in Uncons ained Unde wa e En i onmen s Mona Lü jens* Ha ald S e nbe g* *Ha enCi y Uni e si y Hambu g, Depa men o Hyd og aphy and Geodesy, Henning-Vosche au-Pla z 1, 20457 Hambu g, Ge many (e-mail: {mona.lue jens,ha ald.s e nbe g}@hcu-hambu g.de) Abs ac : Assessing and moni o ing ben hic communi ies is inc easingly impo an in iew o global al e a ion o ma ine en i onmen s. Deep lea ning has p o en o e ec i ely de ec ma ine specimen in unde wa e image y bu s ill ace p oblems wi h small inpu da ase s, uncons ained en i onmen s and class imbalance. This s udy e alua es a da a augmen a ion s a egy o alle ia e hese limi a ions. Th ough syn he ically de i ed image composi ions, he en i e inpu da ase was g ea ly ex ended om 700 o 12700 images. Addi ionally, specimen numbe s o b i le s a s, so co als and glass sponges a e equalized esul ing in a mean a e age p ecision inc ease o 24 %. The o e all mean a e age p ecision o box de ec ions yields 76.7 and o ins ance segmen a ion 67.7 a an in e sec ion o e union h eshold o 0.5. This s udy shows ha deep a chi ec u es such as he deployed Cen e Mask ia ResNeX -101 model can success ully be ained wi h ew o iginal images om a ying unde wa e scenes. Keywo ds: objec de ec ion, deep lea ning, da a augmen a ion, ma ine image y, ben hic mega auna 1. INTRODUCTION Global al e a ion o ma ine en i onmen s due o o e ishing, pollu ion, physical habi a des uc ion and clima e change ha e led o an inc easing decline o animal di e si y and abundance h oughou many ma ine ecosys ems (Jackson, 2008). Especially ben hic mega auna in he Sou he n Ocean a e a isk o en i onmen al change and a e o signi ican ecological alue as hey al e small-scale opog aphy o seabed habi a s a ec ing he en i e ben hic communi y (Gili e al., 2006). Assessing he biodi e si y and cha ac e isa ion o ben hic communi ies is inc easingly impo an o iden i ying ulne able ma ine ecosys ems and de eloping conse a ion s a egies. Ma ine habi a s ha e been s udied based on mainly h ee p ac ices: physical sampling me hods using sledges o awling, acous ical echniques and op ical sys ems. While physical me hods a e able o collec specimen on a lowe axonomical scale hey des uc he en i onmen and hei sampling a e is a he low. Op ical sys ems a e mo e cos e ec i e, obus and p ecise han acous ic sys ems which led o a la ge g owing lib a y o digi al unde wa e images h oughou ecen yea s pa ing he way o new esea ch in au oma ic analy ical me hods. Objec de ec ion, classi ica ion and segmen a ion o image y has become a subs an ial ask in he ield o compu e ision. P io adi ional ea u e desc ip o s a e o en used o ex ac colou , shape o ex u e in o ma ion and a e good a de ec ing speci ic single objec s such as scallops (Dawkins e al., 2013) o lobs e s using classi ie s such as suppo ec o machine (Tan e al., 2018). Howe e , hey a e no obus o a ying unde wa e scenes ha a e exposed o ma ine snow, wa e u bidi y, lens dis o ion, spa se, uns able illumina ion and colou shi due o he su ey pla o m’s a ia ion in speed, angle o al i ude. Mo eo e , he e is a high a iabili y and in a iabili y be ween ea u es belonging o he same class and di e en classes, espec i ely (Pa oni e al., 2021). Deep lea ning (DL) using con olu ional neu al ne wo ks ha e p o en o ou pe o m adi ional based objec de ec ion (Gonzalez-Cid e al., 2017) as hey a e mo e in a ian o he de o ma ion o images. Addi ionally, images can di ec ly be used as inpu wi hou he necessi y o p e-p ocessing. Be e esul s a e o en achie ed using deepe laye s hus inc easing he numbe o pa ame e s o se e al millions. To achie e high pe o mances wi h hose models using images o uncons ained unde wa e scenes o ac oss a ying pla o ms, a la ge aining da ase size is he mos c ucial pa (Langenkämpe e al., 2020), o en, howe e , e y ime- consuming and cos ly o es ablish. This pape in es iga es he e ec o inpu image da a augmen a ion and composi ion s a egies in an a emp o o e come limi a ions o la ge da a se size and class imbalance. Taking accoun o he esul s, he s a e-o - he a ancho - ee objec de ec ion and ins ance segmen a ion model Cen e Mask (Lee and Pa k, 2019) ia ResNeX -101 (Xie e al., 2017) will be ained based on a small, highly di e se 700 image da ase . In his wo k, he de ec ion, segmen a ion and coun ing o glass sponges (hexac inellids), so co als (p imnoids and ch ysogo giids) and b i le s a s (ophiu oids) will be assessed p o iding i s s eps owa ds u u e abundance and size es ima ions o hese specimen. 2. RELATED RESEARCH Se e al p e ious wo ks add ess objec de ec ion and classi ica ion o ma ine scenes using cu ing-edge DL a chi ec u es such as Re inaNe wi h ResNe -50 (Boulais e al., 2020) o FDCNe (Lu e al., 2018) showing good classi ica ion esul s on single-labelled o iconic images. Au oma ic segmen a ion o ben hic auna has been s udied o co als using DeepLab 3+ (Pa oni e al., 2021) and scale wo ms using U-Ne and VGG-16 CNN (Shashidha a e al., 2020). Secu ing enough aining da a o DL is c ucial as desc ibed abo e hence da a augmen a ion echniques a e widely applied. F equen ly used echniques include he change o ligh in ensi y, sha pness, noise and blu ing (Salman e al., 2016) o change o pe spec i e (Huang e al., 2019). Also, o a ion and c opping o unde wa e images ha e been used (Langenkämpe e al., 2020). These me hods ha e p o en o be use ul, howe e , only limi ed image manipula ions can be pe o med. Mos no ably, he numbe o o iginal anno a ed images om mos s a ed wo ks exceed he amoun o a ailable aining da a o his esea ch. The e o e, ano he echnique will be used which changes he en i e image composi ion and adds addi ional al e a ion o syn he ically gene a ed images. Also, no esea ch ega ding ins ance segmen a ion and coun ing o he selec ed mo pho ypes is known o he au ho s. 3. MATERIALS AND METHODS 3.1 Unde wa e Image Da ase The image da ase used o his s udy was collec ed du ing he expedi ion PS118 o he esea ch essel RV Pola s e n in 2019 (Pu se e al., 2021). Images we e sampled using he owed Ocean Floo Obse a ion and Ba hyme y Sys em (Pu se e al., 2019) wi h a lying al i ude o app oxima ely 1.5 – 2.5 m abo e he sea loo . Se en di e en sampling s a ions om he wes e n Weddell Sea con inen al shel o he no he n Powell Basin we e selec ed. Each s a ion a ea ea u es di e en subs a e ypes anging om so and ine mud sedimen o pebbles and complex ocky opog aphy. The o iginal 3840 x 5760 sized images we e iled o 1440 x 960 o main ain esolu ion bu educe he need o compu ing powe . 1000 images we e selec ed and anno a ed using he web-based image segmen a ion ool COCO Anno a o (B ooks, 2019). O he 1000 images, 700 we e used as aining se , 100 as alida ion se and 200 as es se . In o al, 3550 anno a ions o he aining se we e made, o which 87 % belong o he class b i le s a s, 8 % o he class glass sponges and 5 % o he class so co als. Fo he es and alida ion se 85 % and 84 % o he anno a ions belong o he class b i le s a s, 10 % and 12 % o he class glass sponges and 5 % and 4 % o he class so co als, espec i ely. I is appa en ha he e is a high class imbalance. 3.2 Da a Augmen a ion To inc ease he numbe o images o aining, he image gene a o COCO Syn h (Kelly, 2019) was u ilised which composes cu ou o eg ound images o objec s o e andom image backg ounds. Fo eg ounds a e andomly al e ed in scale, amoun , o a ion and b igh ness o each composi ion (Figu e 1). Fo his s udy, 30 o eg ounds pe class and 30 backg ounds om o iginal images we e used o aining. Images o he composi ions we e selec ed om a ying s a ions and di e om p e ious ones in 3.1. O e all, 12,000 syn he ic images we e deployed o aining o which 2000 images we e gene a ed o glass sponges and so co als each, o educe he e ec o class imbalance. Now, 33 % o he o al anno a ions belong o he class glass sponges, 33 % o he class so co als and 34 % o he class b i le s a s. In o de o emphasize he selec ed augmen a ion me hod, se e al equen ly used image manipula ion echniques we e addi ionally pe o med and compa ed o he selec ed me hod. The ollowing image a ibu es we e al e ed: b igh ness, colou one, con as , sa u a ion, sha pness and blu (Figu e 1). In o al, en di e en al e a ions pe image we e conduc ed. Fig. 1. Example images o syn he ically de i ed image composi ions (1s ow) and adi ional da a augmen a ion me hods (2nd ow, om le o igh ): o iginal image, blu , low b igh ness, high b igh ness, blue colou , g een colou , high con as , low con as , high sa u a ion, low sa u a ion, sha pness 78 Mona Lü jens e al. / IFAC Pape sOnLine 54-16 (2021) 76–82 3.3 Deep Lea ning A chi ec u e and T aining The deep lea ning a chi ec u e chosen o his s udy is he ancho - ee one s age ins ance segmen a ion and objec de ec o Cen e Mask (Lee and Pa k, 2019) in combina ion wi h he backbone ne wo k ResNeX -101 (Xie e al., 2017). While objec de ec ion is he cen al ask o abundance and assemblage s udies, p edic ing masks will be impo an o u u e size es ima ions and biomass p edic ions o ben hic species. An ins ance segmen a ion model was he e o e chosen in o de o adap o di e se esea ch ques ions o u u e analyses. Also, he model should ope a e on high in e ence speed while main aining s ong pe o mances o di ec ly e alua e da ase s on boa d o esea ch essels o s a egic da a collec ion. Since bo h selec ed a chi ec u es mee he men ioned equi emen s and u he p oduce excellen esul s in ecen benchma k challenges such as COCO (Lin e al., 2014), hey a e an app op ia e choice o he espec i e compu e ision ask o ma ine imaging. The backbone ResNeX -101 used o ea u e ex ac ion is an ad ancemen o he deep esidual ne wo k ResNe (He e al., 2016) ha has been ecen ly p oposed. Based on ResNe , ResNeX -101 ollows he s a egy o epea ing laye s bu s acks hem pa allel a he han sequen ially. Thus, esul ing in accu acy imp o emen s while educing he ne wo k complexi y and numbe o pa ame e s. Fo his s udy he 101 laye ed ne wo k was used. As a one s age de ec o , Cen e Mask does no ha e a p oposal s ep and p io i izes in e ence speed. Addi ionally, as an ancho - ee de ec o , i does no use p ede ined bounding boxes o iden i y objec s and is hus insensi i e o di e en da ase s and hype -pa ame e s (e.g. inpu size, scales, e c.). Hence, ancho - ee de ec o s alle ia e limi a ions o objec s ha ha e la ge shape a ia ions and a e a he small (Tian e al., 2019) which is ideal o he selec ed classes used in his esea ch. Cen e Mask adop s FCOS (Tian e al., 2019) as de ec ion head ha di ec ly compu es a 4D ec o and a class label a each p oposed loca ion o di e en le els o ea u e maps. Then, he spa ial a en ion-guided mask (SAG-Mask) compu es he segmen a ion masks on each p edic ed box egion using he spa ial a en ion module (SAM) ha helps he mask o ocus on signi ican pixels (Lee and Pa k, 2019). T aining was execu ed on a 64-bi Linux machine equipped wi h an In el® Xeon® Gold 6254 CPU @ 3.10 GHz and 5 NVIDIA® Tesla® V100 GPU. The base lea ning a e was se o 0.002 and educed by a ac o o 10 a e 25400 and again a e 38100 i e a ions. To educe ea ly o e i ing on highly di e en ia ed da ase s, he lea ning a e was also educed o he i s 5080 i e a ions. Addi ionally, a weigh decay was implemen ed. The maximum numbe o i e a ions was 50800 which co esponds o 20 epochs. All backbone models a e ini ialized by ImageNe p e- ained weigh s. 3.4 E alua ion P o ocol To assess he pe o mance o he model, he e alua ion me ics a e age p ecision, a e age ecall, F1 measu e and accu acy a e u ilized. While he p ecision P e lec s he p opo ion o alse posi i es FP, he ecall R de ines he p opo ion o alse nega i es FN and can be ma hema ically exp essed as ollows: P = TP (TP + FP) and R = TP (TP + FN) , (1 ) wi h TP being he numbe o ue posi i e p edic ions. In o de o classi y whe he a p edic ion is a TP o FP, he in e sec ion o e union (IoU) h eshold is used as i measu es he o e lap be ween he g ound u h and he p edic ed bounding box o segmen a ion mask, espec i ely. Typical alues o IoU h esholds a e 0.5 o 0.75. The e y common app oach o summa ize p ecision and ecall in o one alue is he a e age p ecision (AP). The AP o a single class is he a e aged p ecision ac oss all ecall le els. Complemen a ily, he a e age ecall sco e (AR) a e ages ecall alues o e all IoU ∈ [0.5, 1.0] o each class. The mean a e age p ecision (mAP) and mean a e age ecall (mAR) ac oss all classes C a e de ined as (Raphael e al., 2020): 𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚 = 1 𝐶𝐶𝐶𝐶�𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑖𝑖𝑖𝑖 𝐶𝐶𝐶𝐶 𝑖𝑖𝑖𝑖=1 and mAR = 1 𝐶𝐶𝐶𝐶�𝑚𝑚𝑚𝑚𝐴𝐴𝐴𝐴𝑖𝑖𝑖𝑖 𝐶𝐶𝐶𝐶 𝑖𝑖𝑖𝑖=1 (2) As he e alua ion me ics used o his esea ch is based on COCO (Lin e al., 2014), i should be no ed ha mAP and mAR sco es a e u he deno ed as AP and AR o simplici y easons. They a e compu ed o e single (0.5) IoU o he a e age o hen IoU le els s a ing om 0.5 o 0.95 in s eps o 0.05 ( he la e is u he deno ed as AP @.50:.95). AP and AR a e also calcula ed o di e en objec scales (small: < 72² pixels, medium: > 72² & < 214² pixels, la ge: > 214² pixels) and o di e en maximum numbe o de ec ions pe image (1, 10, 100). Objec scales de ia e om COCO and a e adjus ed o i scales in he p oposed images. Addi ional adop ed pe o mance me ics a e he accu acy o assess he o al numbe o p edic ions ha a e co ec and he F1 measu e which e enly weighs be ween p ecision and ecall (Manning e al., 2009): accu acy = TP + TN (TP + TN + FP + FN) 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 𝐹𝐹𝐹𝐹1= 2𝑚𝑚𝑚𝑚𝐴𝐴𝐴𝐴 𝑚𝑚𝑚𝑚+𝐴𝐴𝐴𝐴 (3) 4. EXPERIMENTAL RESULTS This sec ion p esen s he pe o mance e alua ion o di e en in es iga ed me hods o de ec , segmen and coun ben hic mega auna ac oss a ying models and da ase s as lis ed in Table 1. All es uns a e pe o med on he 200 o iginal image es se wi h no da a augmen a ion. Table 1. Abb e ia ions o aining s a egies CM-X-101 Cen e Mask and ResNeX -101 CM-V-99 Cen e Mask and VoVNe V2-99 CM-L-M Cen e Mask-Li e and MobileNe V2 M-X-101 Mask R-CNN and ResNeX -101 R-X-101 Re inaNe and ResNeX -101 Mona Lü jens e al. / IFAC Pape sOnLine 54-16 (2021) 76–82 79 3.3 Deep Lea ning A chi ec u e and T aining The deep lea ning a chi ec u e chosen o his s udy is he ancho - ee one s age ins ance segmen a ion and objec de ec o Cen e Mask (Lee and Pa k, 2019) in combina ion wi h he backbone ne wo k ResNeX -101 (Xie e al., 2017). While objec de ec ion is he cen al ask o abundance and assemblage s udies, p edic ing masks will be impo an o u u e size es ima ions and biomass p edic ions o ben hic species. An ins ance segmen a ion model was he e o e chosen in o de o adap o di e se esea ch ques ions o u u e analyses. Also, he model should ope a e on high in e ence speed while main aining s ong pe o mances o di ec ly e alua e da ase s on boa d o esea ch essels o s a egic da a collec ion. Since bo h selec ed a chi ec u es mee he men ioned equi emen s and u he p oduce excellen esul s in ecen benchma k challenges such as COCO (Lin e al., 2014), hey a e an app op ia e choice o he espec i e compu e ision ask o ma ine imaging. The backbone ResNeX -101 used o ea u e ex ac ion is an ad ancemen o he deep esidual ne wo k ResNe (He e al., 2016) ha has been ecen ly p oposed. Based on ResNe , ResNeX -101 ollows he s a egy o epea ing laye s bu s acks hem pa allel a he han sequen ially. Thus, esul ing in accu acy imp o emen s while educing he ne wo k complexi y and numbe o pa ame e s. Fo his s udy he 101 laye ed ne wo k was used. As a one s age de ec o , Cen e Mask does no ha e a p oposal s ep and p io i izes in e ence speed. Addi ionally, as an ancho - ee de ec o , i does no use p ede ined bounding boxes o iden i y objec s and is hus insensi i e o di e en da ase s and hype -pa ame e s (e.g. inpu size, scales, e c.). Hence, ancho - ee de ec o s alle ia e limi a ions o objec s ha ha e la ge shape a ia ions and a e a he small (Tian e al., 2019) which is ideal o he selec ed classes used in his esea ch. Cen e Mask adop s FCOS (Tian e al., 2019) as de ec ion head ha di ec ly compu es a 4D ec o and a class label a each p oposed loca ion o di e en le els o ea u e maps. Then, he spa ial a en ion-guided mask (SAG-Mask) compu es he segmen a ion masks on each p edic ed box egion using he spa ial a en ion module (SAM) ha helps he mask o ocus on signi ican pixels (Lee and Pa k, 2019). T aining was execu ed on a 64-bi Linux machine equipped wi h an In el® Xeon® Gold 6254 CPU @ 3.10 GHz and 5 NVIDIA® Tesla® V100 GPU. The base lea ning a e was se o 0.002 and educed by a ac o o 10 a e 25400 and again a e 38100 i e a ions. To educe ea ly o e i ing on highly di e en ia ed da ase s, he lea ning a e was also educed o he i s 5080 i e a ions. Addi ionally, a weigh decay was implemen ed. The maximum numbe o i e a ions was 50800 which co esponds o 20 epochs. All backbone models a e ini ialized by ImageNe p e- ained weigh s. 3.4 E alua ion P o ocol To assess he pe o mance o he model, he e alua ion me ics a e age p ecision, a e age ecall, F1 measu e and accu acy a e u ilized. While he p ecision P e lec s he p opo ion o alse posi i es FP, he ecall R de ines he p opo ion o alse nega i es FN and can be ma hema ically exp essed as ollows: P = TP (TP + FP) and R = TP (TP + FN) , (1 ) wi h TP being he numbe o ue posi i e p edic ions. In o de o classi y whe he a p edic ion is a TP o FP, he in e sec ion o e union (IoU) h eshold is used as i measu es he o e lap be ween he g ound u h and he p edic ed bounding box o segmen a ion mask, espec i ely. Typical alues o IoU h esholds a e 0.5 o 0.75. The e y common app oach o summa ize p ecision and ecall in o one alue is he a e age p ecision (AP). The AP o a single class is he a e aged p ecision ac oss all ecall le els. Complemen a ily, he a e age ecall sco e (AR) a e ages ecall alues o e all IoU ∈ [0.5, 1.0] o each class. The mean a e age p ecision (mAP) and mean a e age ecall (mAR) ac oss all classes C a e de ined as (Raphael e al., 2020): 𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚 = 1 𝐶𝐶𝐶𝐶�𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑖𝑖𝑖𝑖 𝐶𝐶𝐶𝐶 𝑖𝑖𝑖𝑖=1 and mAR = 1 𝐶𝐶𝐶𝐶�𝑚𝑚𝑚𝑚𝐴𝐴𝐴𝐴𝑖𝑖𝑖𝑖 𝐶𝐶𝐶𝐶 𝑖𝑖𝑖𝑖=1 (2) As he e alua ion me ics used o his esea ch is based on COCO (Lin e al., 2014), i should be no ed ha mAP and mAR sco es a e u he deno ed as AP and AR o simplici y easons. They a e compu ed o e single (0.5) IoU o he a e age o hen IoU le els s a ing om 0.5 o 0.95 in s eps o 0.05 ( he la e is u he deno ed as AP @.50:.95). AP and AR a e also calcula ed o di e en objec scales (small: < 72² pixels, medium: > 72² & < 214² pixels, la ge: > 214² pixels) and o di e en maximum numbe o de ec ions pe image (1, 10, 100). Objec scales de ia e om COCO and a e adjus ed o i scales in he p oposed images. Addi ional adop ed pe o mance me ics a e he accu acy o assess he o al numbe o p edic ions ha a e co ec and he F1 measu e which e enly weighs be ween p ecision and ecall (Manning e al., 2009): accu acy = TP + TN (TP + TN + FP + FN) 𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎𝑎 𝐹𝐹𝐹𝐹1= 2𝑚𝑚𝑚𝑚𝐴𝐴𝐴𝐴 𝑚𝑚𝑚𝑚+𝐴𝐴𝐴𝐴 (3) 4. EXPERIMENTAL RESULTS This sec ion p esen s he pe o mance e alua ion o di e en in es iga ed me hods o de ec , segmen and coun ben hic mega auna ac oss a ying models and da ase s as lis ed in Table 1. All es uns a e pe o med on he 200 o iginal image es se wi h no da a augmen a ion. Table 1. Abb e ia ions o aining s a egies CM-X-101 Cen e Mask and ResNeX -101 CM-V-99 Cen e Mask and VoVNe V2-99 CM-L-M Cen e Mask-Li e and MobileNe V2 M-X-101 Mask R-CNN and ResNeX -101 R-X-101 Re inaNe and ResNeX -101 Baseline O iginal da ase Syn h Syn he ic image composi ion wi h equal class dis ibu ion & Baseline Syn h-Blncd Syn he ic image composi ion including ex a glass sponges and so co als & Baseline (balanced aining da a se ) T ad. Augm. T adi ional augmen a ion using b igh ness (high , low), colou one (g een, blue), con as ( high, low), sa u a ion (high, low), sha pness, blu & Baseline Fusion Syn h-Blncd wi h 6,000 images and T ad. Augm. wi h 7,000 images combined & Baseline 4.1 E alua ion o de ec ion, segmen a ion and coun ing Fo u u e ben hic assemblage and abundance analyses, he p ecision sco e as well as ecall a e bo h equally impo an as specimen should nei he be misclassi ied no missed. As can be seen in Table 2, he ne wo k CM-X-101 ained on Syn h- Blncd shows he highes bounding box esul s o 76.7 % AP @.50 and 59 % AR @10 compa ed o all o he me hods. De ec ions a e accu a e on a ious backg ounds, illumina ion, came a dis ances and dis o ions (Figu e 2) con i ming ha au oma ic de ec ion and classi ica ion me hods a e possible e en on small highly di e se inpu da ase s. Also, high AP and AR sco es on segmen a ion masks a e achie ed, yielding 67.7 % @.50 and 49.1 % @10, espec i ely (Table 2). Pe o mance o bounding boxes a e sligh ly highe han ins ance segmen a ion masks because e y coa se objec bounda ies a e d awn on each objec including also many i ele an pixels. Ins ance segmen a ion assigns only objec ele an pixels o a label and is he e o e compu a ionally mo e ad anced. I can be u he no ed, ha he ecall a e o single objec s pe image is much lowe han he ecall o mul iple objec s pe image. This accoun s o bo h segmen a ion masks and bounding boxes. Addi ionally, smalle objec s <72² each poo e esul s han la ge objec s which migh be caused by he downsampling in he ResNeX backbone esul ing in ewe ea u es being ex ac ed. Ano he ac o is he ela i ely la ge a io be ween pixel size and objec size o small objec s which migh o en lead o posi ioning e o s when compu ing IoU. The pe o mance o he model on di e en classes can be u he in es iga ed wi h a con usion ma ix (Table 3) which is compu ed using a lowe IoU h eshold o a ou a high ecall especially o small objec s. A i s , i can be seen ha no specimen a e w ongly classi ied be ween classes. B i le s a s ha e he highes p ecision o 99 % and so co als ha e he highes ecall o 93 %. On he o he hand, glass sponges and b i le s a s a e o en no de ec ed (high FN) and glass sponges and so co als a e o en misclassi ied wi h backg ound clu e (high FP). Using he alues o TP, FP and FN, he accu acy amoun s o 87 %, 73 % and 58 % o b i le s a s, so co als and glass sponges, espec i ely, esul ing in a mean accu acy o 73 %. In o al, 76 ou o 82 glass sponges we e coun ed, 49 ou o 41 so co als and 597 ou o 673 b i le s a s leading o a pe cen age a ia ion o -7 %, 20 % and -11 % o each espec i e class. 4.2 E alua ion o da a augmen a ion s a egies The de ec ion pe o mance o ma ine o ganisms imp o es in bo h cases using ei he he da a augmen a ion s a egy o syn he ic image composi ions o he adi ional image manipula ion echniques. In ac , wi h he Syn h-Blncd da ase , he bounding box AP @.50:.95 esul was he mos imp o ed wi h an inc ease o 24 % o e he Baseline da ase . Gene ally, AP and AR on bounding boxes a e sligh ly highe on syn he ically gene a ed image composi ions han adi ional augmen a ion me hods (Table 2). The disad an age o he la e is ha no as many images can be c ea ed wi hou he isk o o e i ing as objec composi ions a e no changing. Syn he ic gene a ed images ha e p o en o be a success ul al e na i e and can be gene a ed in a sho amoun o ime as well. Howe e , wi h ega ds o ins ance segmen a ion pe o mance, syn he ic image da ase s show in e io esul s on AP sco es especially o small and medium sized objec s (Table 2). The eason o his migh be he inco ec downsizing o o eg ounds as hey could appea pixela ed in he p ocess o image c ea ion. Conside ing he p oblem o class imbalance, syn he ic de i ed images ha e he po en ial o easily balance ou numbe s o specimen be ween classes. Table 4 shows ha he de ec ion pe o mance o glass sponges and so co als a e inc eased by 5-16 % o e he Syn h da ase . The pe o mance o b i le s a s, howe e , is no imp o ed. Fig. 2. Example images o de ec ion and segmen a ion esul s o es se wi h CM-X-101/Syn h-Blncd a a ying s a ions. 80 Mona Lü jens e al. / IFAC Pape sOnLine 54-16 (2021) 76–82 4.3 Compa ison wi h o he s a e-o - he-a algo i hms The selec ed Cen e Mask ia ResNeX -101 a chi ec u e was u he compa ed o o he s a e-o - he a de ec o s and backbones such as Re inaNe (Lin e al., 2017), Mask R-CNN (He e al., 2017), MobileNe V2 (Sandle e al., 2018) o VoVNe V2-99 (Lee and Pa k, 2019). Resul s a e demons a ed in Table 2 o bounding boxes and segmen a ion masks. Wi h ega ds o compe ing de ec o s, bo h Cen e Mask and Re inaNe a e one s age de ec o s whe eas Mask R-CNN is a wo s age de ec o ha u ilizes a egion-o -in e es p oposal s ep which ypically p io i izes de ec ion accu acy o e in e ence speed. Mo eo e , Re inaNe and Mask R-CNN use ancho boxes o ea u e de ec ion. F om he esul s in Table 2 i is e iden ha Re inaNe pe o ms nea ly as well as Cen e Mask whe eas Mask R-CNN shows a much lowe pe o mance conside ing bounding boxes and segmen a ion masks. I can be no ed ha one s age ancho - ee de ec o s can pe o m jus as well on unde wa e image y as o he de ec o ypes. Fu he mo e, e alua ions on h ee di e en ly deep backbones we e conduc ed: ResNeX -101 wi h 114.3 million pa ame e s, VoVNe -99 wi h 96 million pa ame e s and MobileNe V2 wi h 28.7 million pa ame e s (Lee and Pa k, 2019). No iceably, MobileNe V2 wi h ewe laye s has conside ably lowe AP and AR esul s han he o he wo. Meanwhile, conside ing bounding boxes, VoVNe -99 pe o ms nea ly as well as ResNeX -101 and wi h ega ds o small objec s, AP and AR esul s a e e en sligh ly highe . The p oblem o de ec ing small objec s is known o inc ease o e y deep backbones as hey need mo e inpu in o ma ion o cope wi h he massi e amoun o pa ame e s (Nguyen e al., 2020). Small objec s a e desc ibed by ewe pixels and migh no be di e se enough o eed he ne wo k wi h su icien in o ma ion inc easing he changes o o e i ing. Table 2. Summa y o de ec ion esul s wi h bounding boxes (1s ow) and segmen a ion masks (2nd ow/cu si e) Table 3. Con usion ma ix o CM-X-101/Syn h-Blncd Table 4. Summa y o pe o mance esul s pe class Model/Da a AP.50:.95 bbox AP.50 bbox APsmall bbox APmedium bbox APla ge bbox AR1 bbox AR10 bbox AR100 bbox ARsmall bbox ARmedium bbox ARla ge bbox CM-X-101/ Baseline 41.7 35.2 68.2 63.0 25.3 5.70 29.3 21.2 54.7 50.7 21.6 19.4 51.6 44.4 55.2 45.2 25.4 8.90 45.1 32.7 70.8 57.9 CM-X-101/ Syn h 48.8 37.3 71.0 63.0 27.4 4.90 39.1 24.1 62.8 51.7 24.7 20.6 58.8 47.9 64.2 49.5 27.9 7.90 57.3 41.2 77.1 59.5 CM-X-101/ Syn h - Blncd 51.8 40.8 76.7 67.7 27.5 4.60 40.2 27.1 66.1 54.8 25.7 22.1 59.0 49.1 63.9 50.6 27.9 7.50 55.7 41.6 77.9 62.5 CM-X-101/ T ad. Augm 48.8 40.2 75.0 70.4 26.9 7.00 38.6 27.6 58.5 52.2 23.0 20.4 55.3 46.8 58.9 47.9 27.2 10.6 50.1 36.8 72.6 58.6 CM-X-101/ Fusion 51.7 41.5 74.1 69.7 27.1 6.60 42.1 29.0 65.1 53.9 24.9 21.7 57.6 47.0 61.6 48.4 27.5 10.6 52.2 37.1 77.0 59.3 CM-V-99/ Syn h 47.9 36.9 72.0 62.9 27.9 4.80 37.0 23.4 62.8 49.8 23.6 20.1 56.6 46.1 61.9 47.8 28.3 7.60 52.6 38.6 77.1 58.2 CM-L-M/ Syn h 27.3 18.4 48.6 32.6 19.1 0.60 19.0 6.30 40.0 34.6 18.3 15.2 39.1 31.8 43.7 33.4 20.0 2.00 34.4 20.3 59.5 50.5 M-X-101/ Syn h 33.3 25.7 53.2 41.6 13.2 1.70 22.7 11.8 53.0 37.9 20.7 15.2 39.2 31.8 40.0 31.9 13.2 3.80 30.6 20.5 60.7 43.6 R-X-101/ Syn h 47.8 70.7 27.9 37.1 62.2 24.2 56.6 61.9 28.4 53.8 76.7 P edic ed Glass Sponges So Co als B i le S a s Backg ound Recall G ound u h Glass Sponges 58 0 0 24 70.1 So Co als 0 38 0 3 92.7 B i le S a s 0 0 591 82 87.8 Backg ound 18 11 6 P ecision 76.3 77.6 99.0 Accu acy: 73 % Glass Sponges So Co als B i le S a s Model/Da a F1 bbox F1 mask AP.50:.95 bbox AP.50:.95 mask F1 bbox F1 mask AP.50:.95 bbox AP.50:.95 mask F1 bbox F1 mask AP.50:.95 bbox AP.50:.95 mask CM-X-101/ Baseline 67.7 67.7 41.4 46.3 70.3 69.5 36.8 45.5 80.2 67.9 46.9 13.9 CM-X-101/ Syn h 67.8 66.2 45.3 46.5 69.8 68.8 48.8 52.6 79.2 65.1 52.4 12.8 CM-X-101/ Syn h-Blncd 71.4 69.2 51.4 54.0 76.8 74.9 51.5 55.9 79.9 64.7 52.6 12.6 Mona Lü jens e al. / IFAC Pape sOnLine 54-16 (2021) 76–82 81 4.3 Compa ison wi h o he s a e-o - he-a algo i hms The selec ed Cen e Mask ia ResNeX -101 a chi ec u e was u he compa ed o o he s a e-o - he a de ec o s and backbones such as Re inaNe (Lin e al., 2017), Mask R-CNN (He e al., 2017), MobileNe V2 (Sandle e al., 2018) o VoVNe V2-99 (Lee and Pa k, 2019). Resul s a e demons a ed in Table 2 o bounding boxes and segmen a ion masks. Wi h ega ds o compe ing de ec o s, bo h Cen e Mask and Re inaNe a e one s age de ec o s whe eas Mask R-CNN is a wo s age de ec o ha u ilizes a egion-o -in e es p oposal s ep which ypically p io i izes de ec ion accu acy o e in e ence speed. Mo eo e , Re inaNe and Mask R-CNN use ancho boxes o ea u e de ec ion. F om he esul s in Table 2 i is e iden ha Re inaNe pe o ms nea ly as well as Cen e Mask whe eas Mask R-CNN shows a much lowe pe o mance conside ing bounding boxes and segmen a ion masks. I can be no ed ha one s age ancho - ee de ec o s can pe o m jus as well on unde wa e image y as o he de ec o ypes. Fu he mo e, e alua ions on h ee di e en ly deep backbones we e conduc ed: ResNeX -101 wi h 114.3 million pa ame e s, VoVNe -99 wi h 96 million pa ame e s and MobileNe V2 wi h 28.7 million pa ame e s (Lee and Pa k, 2019). No iceably, MobileNe V2 wi h ewe laye s has conside ably lowe AP and AR esul s han he o he wo. Meanwhile, conside ing bounding boxes, VoVNe -99 pe o ms nea ly as well as ResNeX -101 and wi h ega ds o small objec s, AP and AR esul s a e e en sligh ly highe . The p oblem o de ec ing small objec s is known o inc ease o e y deep backbones as hey need mo e inpu in o ma ion o cope wi h he massi e amoun o pa ame e s (Nguyen e al., 2020). Small objec s a e desc ibed by ewe pixels and migh no be di e se enough o eed he ne wo k wi h su icien in o ma ion inc easing he changes o o e i ing. Table 2. Summa y o de ec ion esul s wi h bounding boxes (1s ow) and segmen a ion masks (2nd ow/cu si e) Table 3. Con usion ma ix o CM-X-101/Syn h-Blncd Table 4. Summa y o pe o mance esul s pe class Model/Da a AP.50:.95 bbox AP.50 bbox APsmall bbox APmedium bbox APla ge bbox AR1 bbox AR10 bbox AR100 bbox ARsmall bbox ARmedium bbox ARla ge bbox CM-X-101/ Baseline 41.7 35.2 68.2 63.0 25.3 5.70 29.3 21.2 54.7 50.7 21.6 19.4 51.6 44.4 55.2 45.2 25.4 8.90 45.1 32.7 70.8 57.9 CM-X-101/ Syn h 48.8 37.3 71.0 63.0 27.4 4.90 39.1 24.1 62.8 51.7 24.7 20.6 58.8 47.9 64.2 49.5 27.9 7.90 57.3 41.2 77.1 59.5 CM-X-101/ Syn h - Blncd 51.8 40.8 76.7 67.7 27.5 4.60 40.2 27.1 66.1 54.8 25.7 22.1 59.0 49.1 63.9 50.6 27.9 7.50 55.7 41.6 77.9 62.5 CM-X-101/ T ad. Augm 48.8 40.2 75.0 70.4 26.9 7.00 38.6 27.6 58.5 52.2 23.0 20.4 55.3 46.8 58.9 47.9 27.2 10.6 50.1 36.8 72.6 58.6 CM-X-101/ Fusion 51.7 41.5 74.1 69.7 27.1 6.60 42.1 29.0 65.1 53.9 24.9 21.7 57.6 47.0 61.6 48.4 27.5 10.6 52.2 37.1 77.0 59.3 CM-V-99/ Syn h 47.9 36.9 72.0 62.9 27.9 4.80 37.0 23.4 62.8 49.8 23.6 20.1 56.6 46.1 61.9 47.8 28.3 7.60 52.6 38.6 77.1 58.2 CM-L-M/ Syn h 27.3 18.4 48.6 32.6 19.1 0.60 19.0 6.30 40.0 34.6 18.3 15.2 39.1 31.8 43.7 33.4 20.0 2.00 34.4 20.3 59.5 50.5 M-X-101/ Syn h 33.3 25.7 53.2 41.6 13.2 1.70 22.7 11.8 53.0 37.9 20.7 15.2 39.2 31.8 40.0 31.9 13.2 3.80 30.6 20.5 60.7 43.6 R-X-101/ Syn h 47.8 70.7 27.9 37.1 62.2 24.2 56.6 61.9 28.4 53.8 76.7 P edic ed Glass Sponges So Co als B i le S a s Backg ound Recall G ound u h Glass Sponges 58 0 0 24 70.1 So Co als 0 38 0 3 92.7 B i le S a s 0 0 591 82 87.8 Backg ound 18 11 6 P ecision 76.3 77.6 99.0 Accu acy: 73 % Glass Sponges So Co als B i le S a s Model/Da a F1 bbox F1 mask AP.50:.95 bbox AP.50:.95 mask F1 bbox F1 mask AP.50:.95 bbox AP.50:.95 mask F1 bbox F1 mask AP.50:.95 bbox AP.50:.95 mask CM-X-101/ Baseline 67.7 67.7 41.4 46.3 70.3 69.5 36.8 45.5 80.2 67.9 46.9 13.9 CM-X-101/ Syn h 67.8 66.2 45.3 46.5 69.8 68.8 48.8 52.6 79.2 65.1 52.4 12.8 CM-X-101/ Syn h-Blncd 71.4 69.2 51.4 54.0 76.8 74.9 51.5 55.9 79.9 64.7 52.6 12.6 5. CONCLUSION AND FUTURE STEPS In conclusion, he used da a augmen a ion s a egy o syn he ically de i ed image composi ions p o ed o be a good al e na i e o equen ly used augmen a ion echniques. P oblems such as class imbalance can u he easily be alle ia ed and boos he pe o mance o unde ep esen ed classes. De ec ion, segmen a ion and coun ing o ben hic mega auna is a ask ha can be sol ed wi h good pe o mance using ew a ian o iginal inpu images using ancho - ee one s age de ec o s. In compa ison wi h o he models, i is e iden ha he de ec ion p oblem o small objec s is a challenge ye o be sol ed. Addi ionally, u u e s eps in ol e he in oduc ion o mo e ben hic mo pho ypes, imp o ed coun ing o a e duplica ions o o e lapping images and he alloca ion o posi ion and wa e dep h o each de ec ed specimen o u u e assemblage s udies. 6. ACKNOWLEDGEMENTS We hank he anno a o s o he images: Ga in DMello, Diana Rubio and Seyed Liales ani. Fu he we hank he cap ain and c ew o RV Pola s e n as well as he scien i ic pa y o he c uise PS118 o hei suppo . Special hanks go o Au un Pu se and Huw G i i hs o hei suppo , he da a collec ion on boa d and he iden i ica ion o ben hic o ganisms. REFERENCES Boulais, O., Woodwa d, B., Schlining, B., Lunds en, L., Ba na d, K., Bell, K. C. and Ka ija, K. (2020). Fa homNe : An unde wa e image aining da abase o ocean explo a ion and disco e y. a Xi :2007.00114 3, Co nell Uni e si y, Compu e Science, Compu e Vision and Pa e n Recogni ion [cs.CV]. B ooks, J. (2019). COCO Anno a o . URL h ps://gi hub.com/ jsb oks/coco-anno a o /. Dawkins, M., S ewa , C., Gallage , S. and Yo k, A. (2013). Au oma ic scallop de ec ion in ben hic en i onmen s. 2013 IEEE Wo kshop on Applica ions o Compu e Vision (WACV), 160-170. doi: 10.1109/WACV.2013.6475014. Gili, J.-M., A n z, W. E., Palanques, A., O ejas, C., Cla ke, A., Day on, P. K., Isla, E., Teixidó, N., Rossi, S. and López- González, P. J. (2006). A unique assemblage o epiben hic sessile suspension eede s wi h a chaic ea u es in he high- An a c ic. Deep Sea Resea ch Pa II: Topical S udies in Oceanog aphy, olume (53), 1029-1052. doi: 10.1016/j.ds 2.2005.10.021. Gonzalez-Cid, Y., Bu gue a, A., Bonin-Fon , F. and Ma amo os, A. (2017). Machine lea ning and deep lea ning s a egies o iden i y Posidonia meadows in unde wa e images. OCEANS 2017 – Abe deen, 1-5. doi: 10.1109/OCEANSE.2017.8084991. He, K., Zhang, X., Ren, S. and Sun, J. (2016). Deep Residual Lea ning o Image Recogni ion. 2016 IEEE Con e ence on Compu e Vision and Pa e n Recogni ion (CVPR), 770-778. doi: 10.1109/CVPR.2016.90. He, K., Gkioxa i, G., Dollá , P. and Gi shick, R. (2017). Mask R-CNN. a Xi :1703.06870 3, Co nell Uni e si y, Compu e Science, Compu e Vision and Pa e n Recogni ion [cs.CV]. Huang, H., Zhou, H., Yang, X., Zhang, L., Qi, L. and Zang, A.-Y. (2019). Fas e R-CNN o ma ine o ganisms de ec ion and ecogni ion using da a augmen a ion. Neu ocompu ing, olume (337), 372-384. doi: 10.1016/j.neucom.2019.01.084. Jackson, J. B. C. (2008). Ecological ex inc ion and e olu ion in he b a e new ocean. P oceedings o he Na ional Academy o Sciences, 11458-11465. doi: 10.1073/pnas.0802812105. Kelly, A. (2019). COCO Syn h. URL h ps://gi hub.com/ akTwel e/cocosyn h. Langenkämpe , D., an Ke elae , R., Pu se , A. and Na kempe , T. W. (2020). Gea -Induced Concep D i in Ma ine Images and I s E ec on Deep Lea ning Classi ica ion. F on ie s in Ma ine Science, olume (7). doi: 10.3389/ ma s.2020.00506. Lee, Y. and Pa k, J. (2019). Cen e Mask : Real-Time Ancho - F ee Ins ance Segmen a ion. a Xi : 1911.06667 6, Co nell Uni e si y, Compu e Science, Compu e Vision and Pa e n Recogni ion [cs.CV]. Lin, T.-Y., Mai e, M., Belongie, S., Bou de , L., Gi shick, R., Hays, J., Pe ona, P., Ramanan, D., Zi nick, C. L. and Dollá , P. (2014). Mic oso COCO: Common Objec s in Con ex . In: Flee D., Pajdla T., Schiele B., Tuy elaa s T. (ed.), Compu e Vision – ECCV 2014, 740-755. Sp inge , Cham. doi:10.1007/978-3-319-10602-1_48. Lin, T.-Y., Goyal, P., Gi shick, R., He, K. and Dollá , P. (2017). Focal Loss o Dense Objec De ec ion. a Xi : 1708.02002 2, Co nell Uni e si y, Compu e Science, Compu e Vision and Pa e n Recogni ion [cs.CV]. Lu, H., Li, Y., Uemu a, T., Ge, Z., Xu, X., He, L., Se ikawa, S. and Kim, H. (2018). FDCNe : il e ing deep con olu ional ne wo k o ma ine o ganism classi ica ion. Mul imedia Tools and Applica ions, olume (77), 21847- 21860. doi: 10.1007/s11042-017-4585-1. Manning, C. D., Ragha an, P. and Schü ze, H. (2008). In oduc ion o in o ma ion e ie al. Camb idge Uni e si y P ess, Camb idge. ISBN: 0521865719. Nguyen, N.-D., Do, T., Ngo, T. D. and Le, D.-D. (2020). An E alua ion o Deep Lea ning Me hods o Small Objec De ec ion. Jou nal o Elec ical and Compu e Enginee ing, olume (2020), 1-18. doi: 10.1155/2020/3189691. Pa oni, G., Co sini, M., Pede sen, N., Pe o ic, V. and Cignoni, P. (2021). Challenges in he deep lea ning-based seman ic segmen a ion o ben hic communi ies om O ho-images. Applied Geoma ics, olume (13), 131-146. doi: 10.1007/s12518-020-00331-6. Pu se , A., Ma con, Y., D eu e , S., Hoge, U., Sablo ny, B. and Hehemann, L. (2019). Ocean Floo Obse a ion and Ba hyme y Sys em (OFOBS): A New Towed Came a/Sona Sys em o Deep-Sea Habi a Su eys. IEEE Jou nal o Oceanic Enginee ing, olume (44), 87- 99. doi: 10.1109/JOE.2018.2794095. Pu se , A., D eu e , S., G i i hs, H., Hehemann, L., Je osch, K., No dhausen, A., Piepenbu g, D., Rich e , C., Sch öde , H. and Do schel, B. (2021). Seabed ideo and s ill images om he no he n Weddell Sea and he wes e n lanks o he Powell Basin. Ea h Sys em Science Da a, olume (13), 609–615. doi: 10.5194/essd-13-609-2021. 82 Mona Lü jens e al. / IFAC Pape sOnLine 54-16 (2021) 76–82 Raphael, A., Dubinsky, Z., Iluz, D. and Ne anyahu, N. S. (2020). Neu al Ne wo k Recogni ion o Ma ine Ben hos and Co als. Di e si y, olume (12). doi: 10.3390/d12010029. Salman, A., Jalal, A., Sha ai , F., Mian, A., Sho is, M., Seage , J. and Ha ey, E. (2016). Fish species classi ica ion in uncons ained unde wa e en i onmen s based on deep lea ning. Limnology and Oceanog aphy: Me hods, olume (14), 570-585. Doi: 10.1002/lom3.10113. Sandle , M., Howa d, A., Zhu, M., Zhmogino , A. and Chen, L.-C. (2018). MobileNe V2: In e ed Residuals and Linea Bo lenecks. 2018 IEEE/CVF Con e ence on Compu e Vision and Pa e n Recogni ion, 4510-4520. doi: 10.1109/CVPR.2018.00474. Shashidha a, B. M., Sco , M. and Ma bu g, A. (2020). Ins ance Segmen a ion o Ben hic Scale Wo ms a a Hyd o he mal Si e. 2020 IEEE Win e Con e ence on Applica ions o Compu e Vision (WACV), 1303-1312. doi: 10.1109/WACV45572.2020.9093574. Tan, C. S., Lau, P. Y., Co eia, P. L. and Campos, A. (2018). Au oma ic analysis o deep-wa e emo ely ope a ed ehicle oo age o es ima ion o No way lobs e abundance. F on ie s o In o ma ion Technology & Elec onic Enginee ing, olume (19), 1042-1055. doi: 10.1631/FITEE.1700720. Tian, Z., Shen, C., Chen, H. and He, T. (2019). FCOS: Fully Con olu ional One-S age Objec De ec ion. IEEE/CVF In e na ional Con e ence on Compu e Vision (ICCV), 9626-9635. doi: 10.1109/ICCV.2019.00972. Xie, S., Gi shick, R., Dolla , P., Tu, Z. and He, K. (2017). Agg ega ed Residual T ans o ma ions o Deep Neu al Ne wo ks. 2017 IEEE Con e ence on Compu e Vision and Pa e n Recogni ion (CVPR), 5987-5995. doi: 10.1109/CVPR.2017.634.