Full text
TPT- Bench Dataset Title: TPT- Bench: A L arge- Scale, L ong- Term andRobot- EgocentricDatasetforBenchmarkingTarget PersonTracking A uth ors: H anjingY e, Y uZ han, W eixiSitu, G uangchengChen, JingwenY u, Z iqiZ hao, K uanqiCai, A rash A joudani, andH ongZ hang* 1 . O verview ThisdocumentservesastheofficialdatareportforTPT- Bench , alarge- scaledatasetdesignedforTarget PersonTracking( TPT) incrowdedandunstructuredenvironments. Thedatasetemphasizeslong- term trackingchallenges, includingfrequentocclusions, targetdisappearance, andre- identification requirementsfrom arobot- egocentricperspective. Thisdatasetsupportstheresearch paper: " TPT- Bench: A L arge- Scale, L ong- Term andRobot- EgocentricDatasetforBenchmarkingTargetPerson Tracking" . 2 . Eth icalStatementand IR B A pproval Thedatacollectioninvolvedhumansubjectsinpublicspaces. ●C onsent: Explicitinformedconsentwasobtainedfrom theprimarysubjects( targetpersons) before datacollection. ●Privacy: N o personallyidentifiableinformation( PII) ofbystanderswasintentionallyrecordedfor identificationpurposes. Thecollectionadh eresto ethicalguidelinesfordatacollectedinpublic spaces. ●IR B A pproval: Thisstudyreceivedethicalapprovalfrom theSouth ernU niversityofScienceand Tech nologyIR B ( A pprovalN umber: # 2 0 2 5 JSJ1 2 8 ) onA pril0 1 , 2 0 2 5 . 3 . DataO riginand A cquisition 3 . 1 C ollectionPlatform Thedatawascollectedusingacustom- designedsensor- equippedpushingcartthatmimicsthekinematics ofadifferential- drivemobilerobot. Thesensorsuiteincludes: ●PanoramicC amera: RICO H ThetaZ 1 ( 3 6 0 Ddualfisheye) , 1 9 2 0 x9 6 0 resolution@ 3 0 H z. ●Depth C amera: Z ED2 Stereo Camera, 6 4 0 x4 8 0 resolution@ 3 0 H z ( providingRG B- DandV isual- InertialO dometry) . ●3 DL iDA R : O usterO S1 - 6 4 , 6 4 x2 0 4 8 resolution@ 1 0 H z.
●IM U : Built- inIM U from theO usterL iDA R. Sensormountingheights( approximate) : ●PanoramicCamera: 7 2 cm ●L iDA R: 5 4 cm ●Stereo Camera: 4 5 cm 3 . 2 Scenarios Datacollectiontookplaceindiverseindoorandoutdoorenvironments, categorizedinto 1 0 high- level scenarios: 1 . Retail& Shopping( e. g. , supermarkets, malls) 2 . F ood& Dining( e. g. , canteens, restaurants) 3 . Corridors& H allways 4 . Roadways& W alkways 5 . V erticalTransport( Elevators/ Escalators) 6 . O utdoorSquares& Plazas 7 . EducationalRooms 8 . TransportH ubs 9 . Entertainment& L eisure 1 0 . L aboratory& O ffice 4 . DatasetC ontentand O rganization Thedatasetcomprises4 8 sequenceswith atotaldurationofapproximately5 . 3 hours. Thedatais organizedinto thefollowingdirectorystructure: TPT- Bench/ ├── panoramic_images/ │ └── < seq_id> / # e. g. , 0 0 0 1 , 0 0 0 2 . . . │ └── < timestamp> . jpg # Equirectangularprojectionimages ├── G Ts/ │ └── < seq_id> . json # G roundTruth annotations ├── rosbags/ │ └── < seq_id> . bag # Raw sensordata( L iDA R, IM U , Images) ├── descriptions/ │ └── < seq_id> . txt # N aturallanguagedescriptionofthesequence ├── quickview_videos/ │ └── < seq_id> . mp4 # L ow- resvisualizationforquickinspection └── evaluation_results/
└── < seq_id> / └── < baseline> . json # O utputfrom baselinetrackers 4 . 1 DataF ormats ●Images: JPG format, equirectangularprojection. ●R aw Data: RO S Bagformat( . bag) , containingsynchronizedstreamsfrom allsensors. ●A nnotations: JSO N format. 5 . A nnotationand M etadataStandards 5 . 1 L abelingProcess Thedatasetemploysasemi- automaticannotationpipelineto ensureefficiencyandaccuracy: 1 . InitialEstimation: State- of- the- art" M O T + Target- ReID" algorithmswereappliedto generate initialtrackingresults. 2 . M anualR efinement: H umanannotatorsreviewedframesat1 5 H z. A nnotatorscorrectedbounding boxes, handledocclusions, andverifiedtargetidentity. 3 . Interpolation: A nnotationswereupsampledto 3 0 H z usinglinearinterpolation. 4 . V erification: A finalpasswasconductedto verifyconsistencyacrosssequences. 5 . 2 M etadataSch ema( JSO N ) Each < seq_id> . jsonfilecontainsframe- levelmetadatakeyedbytimestamp: { " < timestamp> " : { " is_exist" : 1 , / / 1 = Targetpresent, 0 = Targetabsent/ fullyoccluded " is_behind_glass" : 0 , / / 1 = Behindglass( affectsdepth) , 0 = N otbehindglass, - 1 = A bsent " bbox" : [u, v, w, h], / / 2 DBoundingBox: [top- left- x, top- left- y, width, heigh t] / / Ifabsent: [0 , 0 , 0 , 0 ] " interpolated" : 0 / / 1 = G eneratedviainterpolation, 0 = H umanannotatedframe } } 5 . 3 SequenceDescriptions Each sequenceincludesa< seq_id> . txtfiledescribing: ●Initialtargetappearance( clothing, footwear, hair) .
●Environmentalconditions( lighting: normal/ dim/ dark/ bright) . ●Events( clothingchanges, specificscenario transitions) . 6 . DataIntegrityand A uth enticity ●V erification: A llannotatedframesunderwentvisualinspectionbytrainedannotatorsto ensureth e boundingboxestightlyfitthevisiblepartsofthetargetperson. ●Synch ronization: A llmulti- modaldatastreamsaretime- synchronizedviatheRobotO perating System ( RO S) recordingframework. ●A uth enticity: Thedatasetrepresentsreal- world, unalteredrecordingenvironments. N o synthetic agentswereadded; allpedestriansarerealhumansinnaturalsettings. 7 . C odeand R eproducibility Thecodeforthebenchmarkevaluation, datavisualization, andprocessingtoolkitsishostedseparatelyto allow forversioningandmaintenance. ●C odeR epository: https: / / github. com/ M edlarTea/ TPT- BEN CH - TO O L S ●R elation: Thecoderepositorycontainsthedev_toolkitrequiredto parsetheJSO N annotationsand processthe. bagfilesfoundinthisdataset. 8 . U sageand L icense 8 . 1 L icense ThisdatasetislicensedundertheC reativeC ommonsA ttribution4 . 0 International( C C BY 4 . 0 ) license. Y ouarefreeto shareandadaptthematerialaslongasyouprovideappropriatecredit.