<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"
     xmlns:content="http://purl.org/rss/1.0/modules/content/"
     xmlns:dc="http://purl.org/dc/elements/1.1/"
     xmlns:dcterms="http://purl.org/dc/terms/"
     xmlns:media="http://search.yahoo.com/mrss/"
     xmlns:atom="http://www.w3.org/2005/Atom"
     xmlns:cf="https://www.futureplc.com/rss/content-flags"
>
    <channel>
                    <atom:link href="https://www.tvtechnology.com/feeds/tag/media-qc" rel="self" type="application/rss+xml" />
                            <title><![CDATA[ Latest from Tv Technology in Media-qc ]]></title>
                <link>https://www.tvtechnology.com/tag/media-qc</link>
        <description><![CDATA[ All the latest media-qc content from the Tv Technology team ]]></description>
                                    <lastBuildDate>Wed, 18 Nov 2020 16:07:13 +0000</lastBuildDate>
                            <language>en</language>
                                <item>
                                                            <title><![CDATA[ AI, ML are Pushing Media QC and Monitoring to the Next Level ]]></title>
                                                                                                <dc:content><![CDATA[ <p>Over the years, the complexity of video preparation and delivery has increased dramatically. First, the industry witnessed the move from tape to file-based workflows, followed by the transition from analog to digital. New formats and standards have also emerged, adding to the complexity of video delivery. </p><p>Aside from these technology transformations, consumer viewing habits are shifting. Today’s viewers prefer OTT media services, with 76% of U.S. households subscribing to OTT services compared with 62% for traditional pay-TV, according to the latest research from <a href="https://www.mediapost.com/publications/article/352280/76-of-us-households-have-ott-services-vs-62.html" target="_blank"><u>Parks Associates</u></a>. As broadcasters deliver a higher volume of content to a wider range of screens and global audiences, additional errors are being introduced into the workflow, potentially affecting video and audio quality.</p><p>Recent advancements in automated media quality control and monitoring systems are helping broadcasters deliver error-free video and audio on every screen. In particular, innovations in machine learning and artificial intelligence are pushing media QC and monitoring to the next level, increasing the accuracy and consistency of certain media tasks, including content classification, content categorization, lip sync checks and more.</p><h2 id="media-qc-and-monitoring-is-evolving">MEDIA QC AND MONITORING IS EVOLVING</h2><p>In the early stages of media QC and monitoring, automated systems were limited to simple tasks, such as checking the correctness of audio/video technical parameters, including resolution, frame rate, bitrate, content structure and container parameters.</p><p>Since then, media QC and monitoring has evolved. Today, broadcasters can check for perceptual errors using computer vision and standard audio processing techniques. These checks include interlace artifacts, defective pixels, dropouts, visual text recognition, compression and ghosting artifacts, loudness and language detection. </p><p>With the rise of ML and its success in completing tasks such as content classification and object detection, the scope of media QC and monitoring has expanded. Now broadcasters are using advanced ML techniques capable of semantically understanding content for the purpose of content moderation, content classification, indexing and description generation. Let’s look at a few of the specific media applications that can be optimized with ML and AI technologies.</p><h2 id="speeding-up-content-compliance-with-ml-xa0">SPEEDING UP CONTENT COMPLIANCE WITH ML </h2><p>Monitoring and altering content in order to conform to different rules and regulations is one application that can greatly benefit from ML. Broadcasters must comply with a wide range of rules and regulations, which can vary from one region to another. </p><p>Traditionally, broadcasters have maintained a pool of human moderators to manually filter content for regulatory compliance. Under a typical manual workflow, content is passed through multiple stages of review. If a review fails at any stage, the content goes back for editing. Manual content QC and monitoring is expensive, time-consuming and inaccurate. With so many global and regional aspects of content moderation, it is almost impossible for humans to carry out the job with 100% accuracy.</p><p>By automating this process, broadcasters can eliminate the limitations of manual content moderation, including the inability for people to memorize a significant number of visual symbols and the possibility for human error. With an automated QC and monitoring workflow, broadcasters can more rapidly and accurately check content for the presence of brand names, hate symbols, alcohol, violence, celebrity faces, vulgar speech captions and religious symbols. </p><figure class="van-image-figure " data-bordeaux-image-check ><div class='image-full-width-wrapper'><div class='image-widthsetter' style="max-width:1280px;"><p class="vanilla-image-block" style="padding-top:54.77%;"><img id="WhHZ7foCGR4xqMjchgrkya" name="Interra Systems_TVTech2.jpeg" alt="Interra Systems AI/ML" src="https://cdn.mos.cms.futurecdn.net/WhHZ7foCGR4xqMjchgrkya.jpeg" mos="" align="middle" fullscreen="1" width="1280" height="701" attribution="" endorsement="" class="expandable"><a href='https://cdn.mos.cms.futurecdn.net/WhHZ7foCGR4xqMjchgrkya.jpeg' target='_blank' class='expand-button icon-expand-image icon' ></a></p></div></div><figcaption itemprop="caption description" class=""><span class="credit" itemprop="copyrightHolder">(Image credit: Interra Systems)</span></figcaption></figure><p>When using an automated system powered by ML, computer vision techniques and computer algorithms, the benefits are even greater. ML-based systems can handle huge and multiple content classification check lists without any major performance limitations, driving broadcast workflow efficiencies. </p><p>However, it’s important to note that while current ML solutions are sophisticated and may be combined to create broader applications, they lack the real-world knowledge and human experience needed to create valid and acceptable outcomes on their own. Human input is still required to confirm the validity of patterns and help machines refine the result. Such human interactions are likely to define ML uses in the media industry for the foreseeable future.</p><h2 id="ensuring-superior-quality-captions-with-ml">ENSURING SUPERIOR QUALITY CAPTIONS WITH ML</h2><p>Checking for the presence and accuracy of captions is another application area where ML has proven to be very effective. ML can be used to automatically generate captions where they are not present in the content, check the alignment between captions and audio, and check the correctness of the captions against the spoken audio. In addition, ML simplifies the identification of speakers in audio, ensuring that the correct punctuations are placed in captions. </p><p>Ultimately, with ML, broadcasters can expedite the caption creation and verification processes for both live and VOD content, while ensuring that when content is delivered in multiple video quality levels within OTT video streams, the captions maintain a high quality.</p><p>Over the last decade, automatic speech recognition engines have achieved extremely high accuracy, up to 85%, via ML. Still, automatic speech engines face several challenges, such as robustness issues in noisy environments, the ability to handle variable accents, problems when multiple speakers are talking at the same time, and difficulty with kids’ voices (due to a lack of data to train ML models).</p><p>Keeping humans in the loop is imperative to resolve these challenges. By combining cutting-edge ML and automatic speech recognition technology with a manual review process, broadcasters can bring increased simplicity and cost savings to the creation, management and delivery of captions for traditional TV and video streaming.</p><h2 id="eliminating-av-lip-sync-issues-with-ml">ELIMINATING AV LIP SYNC ISSUES WITH ML</h2><p>Synchronization between audio and video is a common issue today. Leveraging image processing and ML technology and deep neural networks, broadcasters can automatically detect audio and video sync errors. ML offers a faster and more precise approach to detecting audio lead and lag issues in media content, compared with the traditional approach of manually checking for lip sync errors. This allows broadcasters to provide a high quality of experience to viewers (QoE).</p><p>Through the power of ML, broadcasters can perform facial detection, facial tracking, lip detection, lip activity detection and speech identification. With an ML-based lip sync solution, typically one module uses video to extract faces and track lip movement. A second module  uses audio to extract audio features and a third ML module matches the movements with the audio features. Using this technique, it is possible to detect even one frame of synchronization issues.</p><h2 id="conclusion">CONCLUSION</h2><p>The amount of content that broadcasters are delivering across the globe is massive. Ensuring a high-quality video experience on every screen is essential if broadcasters want to keep viewers satisfied. With automated QC and monitoring solutions featuring ML and AI technology, broadcasters are better placed to quickly and more accurately comply with industry and government regulations, deliver high-quality captions, classify and categorize content and eliminate lip sync issues. </p><p><em>Anupama Anantharaman is vice president, Product Management, at Interra Systems.</em></p> ]]></dc:content>
                                                                                                                                            <link>https://www.tvtechnology.com/opinion/ai-ml-are-pushing-media-qc-and-monitoring-to-the-next-level</link>
                                                                            <description>
                            <![CDATA[ Ensuring a high-quality video experience on every screen is essential if broadcasters want to keep viewers satisfied ]]>
                                                                                                            </description>
                                                                                                                                <guid isPermaLink="false">eoxAH2xL3bNfj7uecJbM3f</guid>
                                                                                                <enclosure url="https://cdn.mos.cms.futurecdn.net/Zbt5jJzH8HUp5aKCotCMYJ-1280-80.jpg" type="image/jpeg" length="0"></enclosure>
                                                                        <pubDate>Wed, 18 Nov 2020 16:07:13 +0000</pubDate>                                                                                                                                <updated>Thu, 19 Nov 2020 16:24:57 +0000</updated>
                                                                                                                                            <category><![CDATA[Opinion]]></category>
                                                    <category><![CDATA[Insights]]></category>
                                                                                                                    <dc:creator><![CDATA[ Anupama Anantharaman ]]></dc:creator>                                                                                                        <dc:description><![CDATA[ null ]]></dc:description>
                                                                                                                                <cf:isSponsored>false</cf:isSponsored>
                <cf:hasAffiliateLinks>false</cf:hasAffiliateLinks>
                <cf:isPaid>false</cf:isPaid>
                                                                                                                                <media:content type="image/jpeg" url="https://cdn.mos.cms.futurecdn.net/Zbt5jJzH8HUp5aKCotCMYJ-1280-80.jpg">
                                                            <media:credit><![CDATA[damircudic/Getty Images]]></media:credit>
                                                                                                                                                                                                                                    <media:description><![CDATA[streaming OTT]]></media:description>                                                            <media:text><![CDATA[streaming OTT]]></media:text>
                                <media:title type="plain"><![CDATA[streaming OTT]]></media:title>
                                                    </media:content>
                                                    <media:thumbnail url="https://cdn.mos.cms.futurecdn.net/Zbt5jJzH8HUp5aKCotCMYJ-1280-80.jpg" />
                                                                                                                                                                    <content:encoded >
                            <![CDATA[
                            <article>
                                <p>Over the years, the complexity of video preparation and delivery has increased dramatically. First, the industry witnessed the move from tape to file-based workflows, followed by the transition from analog to digital. New formats and standards have also emerged, adding to the complexity of video delivery. </p><p>Aside from these technology transformations, consumer viewing habits are shifting. Today’s viewers prefer OTT media services, with 76% of U.S. households subscribing to OTT services compared with 62% for traditional pay-TV, according to the latest research from <a href="https://www.mediapost.com/publications/article/352280/76-of-us-households-have-ott-services-vs-62.html" target="_blank"><u>Parks Associates</u></a>. As broadcasters deliver a higher volume of content to a wider range of screens and global audiences, additional errors are being introduced into the workflow, potentially affecting video and audio quality.</p><p>Recent advancements in automated media quality control and monitoring systems are helping broadcasters deliver error-free video and audio on every screen. In particular, innovations in machine learning and artificial intelligence are pushing media QC and monitoring to the next level, increasing the accuracy and consistency of certain media tasks, including content classification, content categorization, lip sync checks and more.</p><h2 id="media-qc-and-monitoring-is-evolving">MEDIA QC AND MONITORING IS EVOLVING</h2><p>In the early stages of media QC and monitoring, automated systems were limited to simple tasks, such as checking the correctness of audio/video technical parameters, including resolution, frame rate, bitrate, content structure and container parameters.</p><p>Since then, media QC and monitoring has evolved. Today, broadcasters can check for perceptual errors using computer vision and standard audio processing techniques. These checks include interlace artifacts, defective pixels, dropouts, visual text recognition, compression and ghosting artifacts, loudness and language detection. </p><p>With the rise of ML and its success in completing tasks such as content classification and object detection, the scope of media QC and monitoring has expanded. Now broadcasters are using advanced ML techniques capable of semantically understanding content for the purpose of content moderation, content classification, indexing and description generation. Let’s look at a few of the specific media applications that can be optimized with ML and AI technologies.</p><h2 id="speeding-up-content-compliance-with-ml-xa0">SPEEDING UP CONTENT COMPLIANCE WITH ML </h2><p>Monitoring and altering content in order to conform to different rules and regulations is one application that can greatly benefit from ML. Broadcasters must comply with a wide range of rules and regulations, which can vary from one region to another. </p><p>Traditionally, broadcasters have maintained a pool of human moderators to manually filter content for regulatory compliance. Under a typical manual workflow, content is passed through multiple stages of review. If a review fails at any stage, the content goes back for editing. Manual content QC and monitoring is expensive, time-consuming and inaccurate. With so many global and regional aspects of content moderation, it is almost impossible for humans to carry out the job with 100% accuracy.</p><p>By automating this process, broadcasters can eliminate the limitations of manual content moderation, including the inability for people to memorize a significant number of visual symbols and the possibility for human error. With an automated QC and monitoring workflow, broadcasters can more rapidly and accurately check content for the presence of brand names, hate symbols, alcohol, violence, celebrity faces, vulgar speech captions and religious symbols. </p><figure class="van-image-figure " data-bordeaux-image-check ><div class='image-full-width-wrapper'><div class='image-widthsetter' style="max-width:1280px;"><p class="vanilla-image-block" style="padding-top:54.77%;"><img id="WhHZ7foCGR4xqMjchgrkya" name="Interra Systems_TVTech2.jpeg" alt="Interra Systems AI/ML" src="https://cdn.mos.cms.futurecdn.net/WhHZ7foCGR4xqMjchgrkya.jpeg" mos="" align="middle" fullscreen="1" width="1280" height="701" attribution="" endorsement="" class="expandable"><a href='https://cdn.mos.cms.futurecdn.net/WhHZ7foCGR4xqMjchgrkya.jpeg' target='_blank' class='expand-button icon-expand-image icon' ></a></p></div></div><figcaption itemprop="caption description" class=""><span class="credit" itemprop="copyrightHolder">(Image credit: Interra Systems)</span></figcaption></figure><p>When using an automated system powered by ML, computer vision techniques and computer algorithms, the benefits are even greater. ML-based systems can handle huge and multiple content classification check lists without any major performance limitations, driving broadcast workflow efficiencies. </p><p>However, it’s important to note that while current ML solutions are sophisticated and may be combined to create broader applications, they lack the real-world knowledge and human experience needed to create valid and acceptable outcomes on their own. Human input is still required to confirm the validity of patterns and help machines refine the result. Such human interactions are likely to define ML uses in the media industry for the foreseeable future.</p><h2 id="ensuring-superior-quality-captions-with-ml">ENSURING SUPERIOR QUALITY CAPTIONS WITH ML</h2><p>Checking for the presence and accuracy of captions is another application area where ML has proven to be very effective. ML can be used to automatically generate captions where they are not present in the content, check the alignment between captions and audio, and check the correctness of the captions against the spoken audio. In addition, ML simplifies the identification of speakers in audio, ensuring that the correct punctuations are placed in captions. </p><p>Ultimately, with ML, broadcasters can expedite the caption creation and verification processes for both live and VOD content, while ensuring that when content is delivered in multiple video quality levels within OTT video streams, the captions maintain a high quality.</p><p>Over the last decade, automatic speech recognition engines have achieved extremely high accuracy, up to 85%, via ML. Still, automatic speech engines face several challenges, such as robustness issues in noisy environments, the ability to handle variable accents, problems when multiple speakers are talking at the same time, and difficulty with kids’ voices (due to a lack of data to train ML models).</p><p>Keeping humans in the loop is imperative to resolve these challenges. By combining cutting-edge ML and automatic speech recognition technology with a manual review process, broadcasters can bring increased simplicity and cost savings to the creation, management and delivery of captions for traditional TV and video streaming.</p><h2 id="eliminating-av-lip-sync-issues-with-ml">ELIMINATING AV LIP SYNC ISSUES WITH ML</h2><p>Synchronization between audio and video is a common issue today. Leveraging image processing and ML technology and deep neural networks, broadcasters can automatically detect audio and video sync errors. ML offers a faster and more precise approach to detecting audio lead and lag issues in media content, compared with the traditional approach of manually checking for lip sync errors. This allows broadcasters to provide a high quality of experience to viewers (QoE).</p><p>Through the power of ML, broadcasters can perform facial detection, facial tracking, lip detection, lip activity detection and speech identification. With an ML-based lip sync solution, typically one module uses video to extract faces and track lip movement. A second module  uses audio to extract audio features and a third ML module matches the movements with the audio features. Using this technique, it is possible to detect even one frame of synchronization issues.</p><h2 id="conclusion">CONCLUSION</h2><p>The amount of content that broadcasters are delivering across the globe is massive. Ensuring a high-quality video experience on every screen is essential if broadcasters want to keep viewers satisfied. With automated QC and monitoring solutions featuring ML and AI technology, broadcasters are better placed to quickly and more accurately comply with industry and government regulations, deliver high-quality captions, classify and categorize content and eliminate lip sync issues. </p><p><em>Anupama Anantharaman is vice president, Product Management, at Interra Systems.</em></p>
                                                            </article>
                            ]]>
                        </content:encoded>
                                                </item>
                                <item>
                                                            <title><![CDATA[ Media Content Compliance is More Efficient, Accurate With ML-Driven QC Solutions ]]></title>
                                                                                                <dc:content><![CDATA[ <p>Machine learning and deep learning-based solutions are making a significant impact on media QC thanks to the availability of large GPU computing power and datasets. Using these technologies, media companies can automatically verify if audiovisual content meets compliance requirements. In regions of the world where nudity, adult content, violence, prohibited objects, substance abuse, and strong language are outlawed, media companies can leverage these technologies to increase compliance reliability and streamline their QC workflows, saving both time and money.</p><p>This article will explain why the success of a learning-based system depends heavily on the quality and quantity of datasets used. While publicly available datasets are good for general development of learning-based systems, they are not adequate for the specific requirements of content compliance in the media industry. Significant efforts are needed to build well-annotated quality datasets for the specific requirements of content compliance. If the training dataset is not well designed, then it is easy for an object detector to confuse guns with cell phones, for example.</p><p><strong>ADVANCEMENTS IN ML/DL TECHNOLOGY</strong></p><p>Content compliance can be a rather intricate process that involves analyzing metadata gathered from a variety of fundamental tasks, such as detecting objects inside frames, recognizing actions over several frames, classifying scenery, detecting specific events in audio or video tracks, classifying videos into specific activities or themes, converting speech to text, and detecting and recognizing faces.</p><p>In a traditional machine learning system, the features extracted from images for content compliance purposes were made by humans. Recent advancements in ML and DL have automated this process. A huge breakthrough in deep learning occurred in 2012 when AlexNet was designed. AlexNet is a convolutional neural network trained on 1.2 million real world images from a dataset called ImageNet for classification purposes. Images are classified into 1000 different categories, five layers and 60 million parameters, making AlexNet one of the most intricate and low-error-rate networks.</p><figure class="van-image-figure pull-" data-bordeaux-image-check ><div class='image-full-width-wrapper'><div class='image-widthsetter' ><p class="vanilla-image-block" style="padding-top:56.25%;"><img id="5qyponWNG6TTpeCSFqFn" name="" alt="Fig. 1: Activity recognition enabled by deep convolutional network." src="https://cdn.mos.cms.futurecdn.net/5qyponWNG6TTpeCSFqFn.jpg" mos="https://cdn.mos.cms.futurecdn.net/5qyponWNG6TTpeCSFqFn.jpg" align="" fullscreen="" width="" height="" attribution="" endorsement="" class="pull-"></p></div></div><figcaption itemprop="caption description" class="pull-"><span class="caption-text">Fig. 1: Activity recognition enabled by deep convolutional network. </span></figcaption></figure><p>After AlexNet there were several additional developments in between the years of 2012 and 2015. Faster R-CNN, a deep neural network for object detection tasks, is one network that was proposed. While AlexNet addresses image classification, Faster R-CNN is designed to resolve object detection problems; therefore, it is more complex since it involves locating the object inside an image. Faster R-CNN recommends possible regions in an image that might contain an object and checks whether the proposed regions contain an object among the list of supported categories or not. If they do, the network returns the bounding box of the region containing the object and the name of object.</p><figure class="van-image-figure pull-" data-bordeaux-image-check ><div class='image-full-width-wrapper'><div class='image-widthsetter' ><p class="vanilla-image-block" style="padding-top:56.25%;"><img id="qoN6ry8XzJiU8859KN9pCk" name="" alt="Figure 2. DL today is based on AlexNet (L) and Faster R-CNN." src="https://cdn.mos.cms.futurecdn.net/qoN6ry8XzJiU8859KN9pCk.jpg" mos="https://cdn.mos.cms.futurecdn.net/qoN6ry8XzJiU8859KN9pCk.jpg" align="" fullscreen="" width="" height="" attribution="" endorsement="" class="pull-"></p></div></div><figcaption itemprop="caption description" class="pull-"><span class="caption-text">Figure 2. DL today is based on AlexNet (L) and Faster R-CNN. </span></figcaption></figure><p>There are two key parts involved with constructing an ML network for QC. First, a QC solutions provider has to train the network on datasets so that the network can start recognizing objects of interest (e.g., guns, alcohol, cigarettes, belly buttons, etc.). Transfer learning is a technique that can be useful when training a network. Transfer learning reuses a trained model as a starting point for training on another dataset. This aids in training a network quickly for new types of objects and with less number of examples. Second, the trained network is applied in the media QC environment to make predictions about the presence of these objects in media files.</p><p>A critical factor of success for deep learning has been the availability of huge well labeled datasets. If datasets are well labeled, they can outperform the accuracy of human visual recognition. In fact, the top-performing dataset model achieved in accuracy of 96 percent in 2017.</p><p><strong>THREE WAYS TO APPLY ML/DL TO MEDIA QC</strong></p><p>ML/DL can be used for a range of different quality and compliance purposes in media workflows. Aside from detecting objects, the technology is useful for recognizing activity in a video frame, onscreen visual text, audio events, and whether captions are aligned correctly. Let’s look at three key ways that operators can use these techniques to their advantage.</p><p>Identifying explicit content is one area where ML/DL technology can be useful in the media environment. Object detection, activity recognition, audio and visual cues can be utilized to determine if there is nudity or minimal covering, mild sexual situations, or explicit sexual situations. Additionally, activity recognition and object detection can be used to identify violence, including the presence of guns, killing, and car crashes. In some regions of the world, the presence of alcohol and smoking in video content is prohibited. Operators can use object detection ML technology to identify alcohol as well as cigarettes, cigars and other vaping devices. Activity recognition plays a role in detecting the actual physical act of smoking.</p><figure class="van-image-figure pull-" data-bordeaux-image-check ><div class='image-full-width-wrapper'><div class='image-widthsetter' ><p class="vanilla-image-block" style="padding-top:56.25%;"><img id="7DT9gEkA7yP5wdWVXsbS6H" name="" alt="Fig. 3: Challenges an object detection system faces during classification" src="https://cdn.mos.cms.futurecdn.net/7DT9gEkA7yP5wdWVXsbS6H.jpg" mos="https://cdn.mos.cms.futurecdn.net/7DT9gEkA7yP5wdWVXsbS6H.jpg" align="" fullscreen="" width="" height="" attribution="" endorsement="" class="pull-"></p></div></div><figcaption itemprop="caption description" class="pull-"><span class="caption-text">Fig. 3: Challenges an object detection system faces during classification </span></figcaption></figure><p><strong>CONCLUSION</strong></p><p>Accuracy is crucial when it comes to quality control for media operations. If an operator misses a video scene with violence or alcohol, they run the risk of not adhering to content compliance requirements. Over the years, learning-based systems have evolved to where datasets are richer and higher in quality, improving the automatic identification of content for compliance purposes. With the latest innovations in ML/DL technology, operators can significantly increase the efficiency and accuracy of their media workflows. </p><p>Interra Systems’ software-based QC solution has been integrated with the latest advancements in ML and AI technology, allowing operators to deliver exceptional audio-video quality on every device and comply with all regional content guidelines.</p><p><em>Shailesh Kumar is Associate Director of Engineering at Interra Systems.</em></p> ]]></dc:content>
                                                                                                                                            <link>https://www.tvtechnology.com/opinions/media-content-compliance-is-more-efficient-accurate-with-ml-driven-qc-solutions</link>
                                                                            <description>
                            <![CDATA[ Success of a learning-based system depends heavily on the quality and quantity of datasets used. ]]>
                                                                                                            </description>
                                                                                                                                <guid isPermaLink="false">iu3sh79quUcNe52k5AkoN7</guid>
                                                                                                <enclosure url="https://cdn.mos.cms.futurecdn.net/7DT9gEkA7yP5wdWVXsbS6H-1280-80.jpg" type="image/jpeg" length="0"></enclosure>
                                                                        <pubDate>Mon, 26 Nov 2018 16:27:21 +0000</pubDate>                                                                                                                                                                                                                                <category><![CDATA[Opinion]]></category>
                                                    <category><![CDATA[Insights]]></category>
                                                                                                                    <dc:creator><![CDATA[ Shailesh Kumar ]]></dc:creator>                                                                                                        <dc:description><![CDATA[ null ]]></dc:description>
                                                                                                                                <cf:isSponsored>false</cf:isSponsored>
                <cf:hasAffiliateLinks>false</cf:hasAffiliateLinks>
                <cf:isPaid>false</cf:isPaid>
                                                                                                                                <media:content type="image/jpeg" url="https://cdn.mos.cms.futurecdn.net/7DT9gEkA7yP5wdWVXsbS6H-1280-80.jpg">
                                                            <media:credit><![CDATA[null]]></media:credit>
                                                                                                                                                                        <media:description><![CDATA[Fig. 3: Challenges an object detection system faces during classification]]></media:description>                                                    </media:content>
                                                    <media:thumbnail url="https://cdn.mos.cms.futurecdn.net/7DT9gEkA7yP5wdWVXsbS6H-1280-80.jpg" />
                                                                                                                                                                    <content:encoded >
                            <![CDATA[
                            <article>
                                <p>Machine learning and deep learning-based solutions are making a significant impact on media QC thanks to the availability of large GPU computing power and datasets. Using these technologies, media companies can automatically verify if audiovisual content meets compliance requirements. In regions of the world where nudity, adult content, violence, prohibited objects, substance abuse, and strong language are outlawed, media companies can leverage these technologies to increase compliance reliability and streamline their QC workflows, saving both time and money.</p><p>This article will explain why the success of a learning-based system depends heavily on the quality and quantity of datasets used. While publicly available datasets are good for general development of learning-based systems, they are not adequate for the specific requirements of content compliance in the media industry. Significant efforts are needed to build well-annotated quality datasets for the specific requirements of content compliance. If the training dataset is not well designed, then it is easy for an object detector to confuse guns with cell phones, for example.</p><p><strong>ADVANCEMENTS IN ML/DL TECHNOLOGY</strong></p><p>Content compliance can be a rather intricate process that involves analyzing metadata gathered from a variety of fundamental tasks, such as detecting objects inside frames, recognizing actions over several frames, classifying scenery, detecting specific events in audio or video tracks, classifying videos into specific activities or themes, converting speech to text, and detecting and recognizing faces.</p><p>In a traditional machine learning system, the features extracted from images for content compliance purposes were made by humans. Recent advancements in ML and DL have automated this process. A huge breakthrough in deep learning occurred in 2012 when AlexNet was designed. AlexNet is a convolutional neural network trained on 1.2 million real world images from a dataset called ImageNet for classification purposes. Images are classified into 1000 different categories, five layers and 60 million parameters, making AlexNet one of the most intricate and low-error-rate networks.</p><figure class="van-image-figure pull-" data-bordeaux-image-check ><div class='image-full-width-wrapper'><div class='image-widthsetter' ><p class="vanilla-image-block" style="padding-top:56.25%;"><img id="5qyponWNG6TTpeCSFqFn" name="" alt="Fig. 1: Activity recognition enabled by deep convolutional network." src="https://cdn.mos.cms.futurecdn.net/5qyponWNG6TTpeCSFqFn.jpg" mos="https://cdn.mos.cms.futurecdn.net/5qyponWNG6TTpeCSFqFn.jpg" align="" fullscreen="" width="" height="" attribution="" endorsement="" class="pull-"></p></div></div><figcaption itemprop="caption description" class="pull-"><span class="caption-text">Fig. 1: Activity recognition enabled by deep convolutional network. </span></figcaption></figure><p>After AlexNet there were several additional developments in between the years of 2012 and 2015. Faster R-CNN, a deep neural network for object detection tasks, is one network that was proposed. While AlexNet addresses image classification, Faster R-CNN is designed to resolve object detection problems; therefore, it is more complex since it involves locating the object inside an image. Faster R-CNN recommends possible regions in an image that might contain an object and checks whether the proposed regions contain an object among the list of supported categories or not. If they do, the network returns the bounding box of the region containing the object and the name of object.</p><figure class="van-image-figure pull-" data-bordeaux-image-check ><div class='image-full-width-wrapper'><div class='image-widthsetter' ><p class="vanilla-image-block" style="padding-top:56.25%;"><img id="qoN6ry8XzJiU8859KN9pCk" name="" alt="Figure 2. DL today is based on AlexNet (L) and Faster R-CNN." src="https://cdn.mos.cms.futurecdn.net/qoN6ry8XzJiU8859KN9pCk.jpg" mos="https://cdn.mos.cms.futurecdn.net/qoN6ry8XzJiU8859KN9pCk.jpg" align="" fullscreen="" width="" height="" attribution="" endorsement="" class="pull-"></p></div></div><figcaption itemprop="caption description" class="pull-"><span class="caption-text">Figure 2. DL today is based on AlexNet (L) and Faster R-CNN. </span></figcaption></figure><p>There are two key parts involved with constructing an ML network for QC. First, a QC solutions provider has to train the network on datasets so that the network can start recognizing objects of interest (e.g., guns, alcohol, cigarettes, belly buttons, etc.). Transfer learning is a technique that can be useful when training a network. Transfer learning reuses a trained model as a starting point for training on another dataset. This aids in training a network quickly for new types of objects and with less number of examples. Second, the trained network is applied in the media QC environment to make predictions about the presence of these objects in media files.</p><p>A critical factor of success for deep learning has been the availability of huge well labeled datasets. If datasets are well labeled, they can outperform the accuracy of human visual recognition. In fact, the top-performing dataset model achieved in accuracy of 96 percent in 2017.</p><p><strong>THREE WAYS TO APPLY ML/DL TO MEDIA QC</strong></p><p>ML/DL can be used for a range of different quality and compliance purposes in media workflows. Aside from detecting objects, the technology is useful for recognizing activity in a video frame, onscreen visual text, audio events, and whether captions are aligned correctly. Let’s look at three key ways that operators can use these techniques to their advantage.</p><p>Identifying explicit content is one area where ML/DL technology can be useful in the media environment. Object detection, activity recognition, audio and visual cues can be utilized to determine if there is nudity or minimal covering, mild sexual situations, or explicit sexual situations. Additionally, activity recognition and object detection can be used to identify violence, including the presence of guns, killing, and car crashes. In some regions of the world, the presence of alcohol and smoking in video content is prohibited. Operators can use object detection ML technology to identify alcohol as well as cigarettes, cigars and other vaping devices. Activity recognition plays a role in detecting the actual physical act of smoking.</p><figure class="van-image-figure pull-" data-bordeaux-image-check ><div class='image-full-width-wrapper'><div class='image-widthsetter' ><p class="vanilla-image-block" style="padding-top:56.25%;"><img id="7DT9gEkA7yP5wdWVXsbS6H" name="" alt="Fig. 3: Challenges an object detection system faces during classification" src="https://cdn.mos.cms.futurecdn.net/7DT9gEkA7yP5wdWVXsbS6H.jpg" mos="https://cdn.mos.cms.futurecdn.net/7DT9gEkA7yP5wdWVXsbS6H.jpg" align="" fullscreen="" width="" height="" attribution="" endorsement="" class="pull-"></p></div></div><figcaption itemprop="caption description" class="pull-"><span class="caption-text">Fig. 3: Challenges an object detection system faces during classification </span></figcaption></figure><p><strong>CONCLUSION</strong></p><p>Accuracy is crucial when it comes to quality control for media operations. If an operator misses a video scene with violence or alcohol, they run the risk of not adhering to content compliance requirements. Over the years, learning-based systems have evolved to where datasets are richer and higher in quality, improving the automatic identification of content for compliance purposes. With the latest innovations in ML/DL technology, operators can significantly increase the efficiency and accuracy of their media workflows. </p><p>Interra Systems’ software-based QC solution has been integrated with the latest advancements in ML and AI technology, allowing operators to deliver exceptional audio-video quality on every device and comply with all regional content guidelines.</p><p><em>Shailesh Kumar is Associate Director of Engineering at Interra Systems.</em></p>
                                                            </article>
                            ]]>
                        </content:encoded>
                                                </item>
            </channel>
</rss>