Matches in Ghent University Academic Bibliography for { <https://biblio.ugent.be/publication/01GMBA3JPV0KX0YXEGDZVFSSMN> ?p ?o. }
Showing items 1 to 24 of
24
with 100 items per page.
- 01GMBA3JPV0KX0YXEGDZVFSSMN classification P1.
- 01GMBA3JPV0KX0YXEGDZVFSSMN date "2022".
- 01GMBA3JPV0KX0YXEGDZVFSSMN language "eng".
- 01GMBA3JPV0KX0YXEGDZVFSSMN type conference.
- 01GMBA3JPV0KX0YXEGDZVFSSMN hasPart 01GMBA8DVTP0GJS01BTKSK8GVC.pdf.
- 01GMBA3JPV0KX0YXEGDZVFSSMN subject "Technology and Engineering".
- 01GMBA3JPV0KX0YXEGDZVFSSMN doi "10.21437/Interspeech.2022-196".
- 01GMBA3JPV0KX0YXEGDZVFSSMN issn "2308-457X".
- 01GMBA3JPV0KX0YXEGDZVFSSMN presentedAt urn:uuid:ad2cfd73-321a-4c31-a272-bfa3d2a36610.
- 01GMBA3JPV0KX0YXEGDZVFSSMN abstract "Sequential audio event tagging can provide not only the type information of audio events, but also the order information between events and the number of events that occur in an audio clip. Most previous works on audio event sequence analysis rely on connectionist temporal classification (CTC). However, CTC's conditional independence assumption prevents it from effectively learning correlations between diverse audio events. This paper first introduces the Transformer into sequential audio tagging, since Transformers perform well in sequence-related tasks. To better utilize contextual information of audio event sequences, we draw on the idea of bidirectional recurrent neural networks, and propose a contextual Transformer (cTransformer) with a bidirectional decoder that could exploit the forward and backward information of event sequences. Experiments on the real-life polyphonic audio dataset show that, compared to CTC-based methods, the cTransformer can effectively combine the fine-grained acoustic representations from the encoder and coarse-grained audio event cues to exploit contextual information to successfully recognize and predict the audio event sequence in polyphonic audio clips.".
- 01GMBA3JPV0KX0YXEGDZVFSSMN author 7E14BF6C-50F5-11E5-B4A0-F149B5D1D7B1.
- 01GMBA3JPV0KX0YXEGDZVFSSMN author F43ABB58-F0ED-11E1-A9DE-61C894A0A6B4.
- 01GMBA3JPV0KX0YXEGDZVFSSMN author bf552237-bd78-11ea-9edd-84a31b5b5824.
- 01GMBA3JPV0KX0YXEGDZVFSSMN author urn:uuid:32d850ab-c6cf-47cf-ab4c-6e3a506a4aa1.
- 01GMBA3JPV0KX0YXEGDZVFSSMN author urn:uuid:56565320-884c-4788-b783-82bddf6d3275.
- 01GMBA3JPV0KX0YXEGDZVFSSMN dateCreated "2022-12-15T16:33:00Z".
- 01GMBA3JPV0KX0YXEGDZVFSSMN dateModified "2024-07-09T07:46:37Z".
- 01GMBA3JPV0KX0YXEGDZVFSSMN name "CT-SAT : contextual transformer for sequential audio tagging".
- 01GMBA3JPV0KX0YXEGDZVFSSMN pagination urn:uuid:4aa27b90-7eff-422a-a839-ea564f417fe2.
- 01GMBA3JPV0KX0YXEGDZVFSSMN publisher urn:uuid:9c9af3ef-d37c-4f85-b90d-129e93436c7d.
- 01GMBA3JPV0KX0YXEGDZVFSSMN sameAs LU-01GMBA3JPV0KX0YXEGDZVFSSMN.
- 01GMBA3JPV0KX0YXEGDZVFSSMN sourceOrganization urn:uuid:25ffaa8d-ab33-402c-9c0b-859d51732703.
- 01GMBA3JPV0KX0YXEGDZVFSSMN sourceOrganization urn:uuid:6f21e54d-4f0e-4886-959e-d5686339e1bc.
- 01GMBA3JPV0KX0YXEGDZVFSSMN type P1.