{"id":82402,"date":"2024-07-25T18:51:51","date_gmt":"2024-07-25T18:51:51","guid":{"rendered":"https:\/\/neclink.com\/index.php\/2024\/07\/25\/ai-video-startup-runway-reportedly-trained-on-thousands-of-youtube-videos-without-permission\/"},"modified":"2024-07-25T18:51:51","modified_gmt":"2024-07-25T18:51:51","slug":"ai-video-startup-runway-reportedly-trained-on-thousands-of-youtube-videos-without-permission","status":"publish","type":"post","link":"https:\/\/neclink.com\/index.php\/2024\/07\/25\/ai-video-startup-runway-reportedly-trained-on-thousands-of-youtube-videos-without-permission\/","title":{"rendered":"AI video startup Runway reportedly trained on \u2018thousands\u2019 of YouTube videos without permission"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p>AI company Runway reportedly scraped \u201cthousands\u201d of YouTube videos and pirated versions of copyrighted movies without permission. <em>404 Media<\/em> <a data-i13n=\"elm:context_link;elmt:doNotAffiliate;cpos:1;pos:1\" class=\"link \" href=\"https:\/\/www.404media.co\/runway-ai-image-generator-training-data-youtube\/\" rel=\"nofollow noopener\" target=\"_blank\" data-ylk=\"slk:obtained;elm:context_link;elmt:doNotAffiliate;cpos:1;pos:1;itc:0;sec:content-canvas\">obtained<\/a> alleged internal spreadsheets suggesting the AI video-generating startup trained its Gen-3 model using YouTube content from channels like Disney, Netflix, Pixar and popular media outlets.<\/p>\n<p>An alleged former Runway employee told the publication the company used the spreadsheet to flag lists of videos it wanted in its database. It would then download them without detection using open-source proxy software to cover its tracks. One sheet lists simple keywords like astronaut, fairy and rainbow, with footnotes indicating whether the company had found corresponding high-quality videos to train on. For example, the term \u201csuperhero\u201d includes a note reading, \u201cLots of movie clips.\u201d (Indeed.)<\/p>\n<p>Other notes show Runway flagged YouTube channels for Unreal Engine, filmmaker Josh Neuman and a Call of Duty fan page as good sources for \u201chigh movement\u201d training videos.<\/p>\n<p>\u201cThe channels in that spreadsheet were a company-wide effort to find good quality videos to build the model with,\u201d the former employee told <em>404 Media<\/em>. \u201cThis was then used as input to a massive web crawler which downloaded all the videos from all those channels, using proxies to avoid getting blocked by Google.\u201d<\/p>\n<figure class=\"caas-figure\">\n<div class=\"caas-figure-with-pb\" style=\"max-height: 540px\">\n<div>\n<div class=\"caas-img-container caas-img-loader\" style=\"padding-bottom:56%\"><img decoding=\"async\" class=\"caas-img caas-lazy has-preview\" alt=\"Screnshot of the Runway AI homepad. \" src=\"https:\/\/s.yimg.com\/ny\/api\/res\/1.2\/972Ki7PD4cV6y0_ZnCps5g--\/YXBwaWQ9aGlnaGxhbmRlcjt3PTk2MDtoPTU0MA--\/https:\/\/s.yimg.com\/os\/creatr-uploaded-images\/2024-07\/0ccba6a0-4aaf-11ef-a7f8-30fbb9c3ef79\"\/><img decoding=\"async\" alt=\"Screnshot of the Runway AI homepad. \" src=\"https:\/\/s.yimg.com\/ny\/api\/res\/1.2\/972Ki7PD4cV6y0_ZnCps5g--\/YXBwaWQ9aGlnaGxhbmRlcjt3PTk2MDtoPTU0MA--\/https:\/\/s.yimg.com\/os\/creatr-uploaded-images\/2024-07\/0ccba6a0-4aaf-11ef-a7f8-30fbb9c3ef79\" class=\"caas-img\"\/><\/div>\n<\/div>\n<\/div>\n<p><figcaption class=\"caption-collapse\"><span class=\"caption-credit\"> Runway<\/span><\/figcaption><\/p>\n<\/figure>\n<p>A list of nearly 4,000 YouTube channels, compiled in one of the spreadsheets, flagged \u201crecommended channels\u201d from CBS New York, AMC Theaters, Pixar, Disney Plus, Disney CD and the Monterey Bay Aquarium. (Because no AI model is complete without otters.)<\/p>\n<p>In addition, Runway reportedly compiled a separate list of videos from piracy sites. A spreadsheet titled \u201cNon-YouTube Source\u201d includes 14 links to sources like an unauthorized online archive of <a data-i13n=\"cpos:2;pos:1\" href=\"https:\/\/www.engadget.com\/studio-ghiblis-the-boy-and-the-heron-arrives-on-max-in-september-140046955.html\" data-ylk=\"slk:Studio Ghibli films;cpos:2;pos:1;elm:context_link;itc:0;sec:content-canvas\" class=\"link \">Studio Ghibli films<\/a>, anime and movie piracy sites, a fan site displaying Xbox game videos and the animated streaming site kisscartoon.sh.<\/p>\n<p>In what could be viewed as a damning confirmation that the company used the training data, <em>404 Media<\/em> found that prompting the video generator with the names of popular YouTubers listed in the spreadsheet spit out results bearing an uncanny resemblance. Crucially, entering the same names in Runway\u2019s older Gen-2 model \u2014 trained before the alleged data in the spreadsheets \u2014\u00a0generated \u201cunrelated\u201d results like generic men in suits. Additionally, after the publication contacted Runway asking about the YouTubers\u2019 likenesses appearing in results, the AI tool stopped generating them altogether.<\/p>\n<p>\u201cI hope that by sharing this information, people will have a better understanding of the scale of these companies and what they\u2019re doing to make \u2018cool\u2019 videos,\u201d the former employee told <em>404 Media<\/em>.<\/p>\n<p>When contacted for comment, a YouTube representative pointed Engadget to an <a data-i13n=\"elm:context_link;elmt:doNotAffiliate;cpos:3;pos:1\" class=\"link \" href=\"https:\/\/www.bloomberg.com\/news\/articles\/2024-04-04\/youtube-says-openai-training-sora-with-its-videos-would-break-the-rules\" rel=\"nofollow noopener\" target=\"_blank\" data-ylk=\"slk:interview;elm:context_link;elmt:doNotAffiliate;cpos:3;pos:1;itc:0;sec:content-canvas\">interview<\/a> its CEO Neal Mohan gave to <em>Bloomberg<\/em> in April. In that interview, Mohan described training on its videos as a \u201cclear violation\u201d of its terms. \u201cOur previous comments on this still stand,\u201d YouTube spokesperson Jack Mason wrote to Engadget.<\/p>\n<p>Runway did not respond to a request for commeInt by the time of publication.<\/p>\n<p>At least some AI companies appear to be in a race to normalize their tools and establish market leadership before users \u2014 and courts \u2014 catch onto how their sausage was made. Training with permission through licensed deals is one thing, and that\u2019s another tactic companies like <a data-i13n=\"cpos:4;pos:1\" href=\"https:\/\/www.engadget.com\/openai-will-train-its-ai-models-on-the-financial-times-journalism-173249177.html\" data-ylk=\"slk:OpenAI have recently adopted;cpos:4;pos:1;elm:context_link;itc:0;sec:content-canvas\" class=\"link \">OpenAI have recently adopted<\/a>. But it\u2019s a much sketchier (if not illegal) proposition to treat the entire internet \u2014 copyrighted material and all \u2014 as up for grabs in a breakneck race for profit and dominance.<\/p>\n<p><em>404 Media<\/em>\u2019s excellent <a data-i13n=\"elm:context_link;elmt:doNotAffiliate;cpos:5;pos:1\" class=\"link \" href=\"https:\/\/www.404media.co\/runway-ai-image-generator-training-data-youtube\/\" rel=\"nofollow noopener\" target=\"_blank\" data-ylk=\"slk:reporting is worth a read;elm:context_link;elmt:doNotAffiliate;cpos:5;pos:1;itc:0;sec:content-canvas\">reporting is worth a read<\/a>.<\/p>\n<\/div>\n<p><br \/>\n<br \/><a href=\"https:\/\/www.engadget.com\/ai-video-startup-runway-reportedly-trained-on-thousands-of-youtube-videos-without-permission-182314160.html?src=rss\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>AI company Runway reportedly scraped \u201cthousands\u201d of YouTube videos and pirated versions of copyrighted movies without permission. 404 Media obtained alleged internal spreadsheets suggesting the<\/p>\n","protected":false},"author":1,"featured_media":82403,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_monsterinsights_skip_tracking":false,"_monsterinsights_sitenote_active":false,"_monsterinsights_sitenote_note":"","_monsterinsights_sitenote_category":0,"footnotes":""},"categories":[157],"tags":[],"class_list":["post-82402","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-gadget"],"aioseo_notices":[],"_links":{"self":[{"href":"https:\/\/neclink.com\/index.php\/wp-json\/wp\/v2\/posts\/82402","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/neclink.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/neclink.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/neclink.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/neclink.com\/index.php\/wp-json\/wp\/v2\/comments?post=82402"}],"version-history":[{"count":0,"href":"https:\/\/neclink.com\/index.php\/wp-json\/wp\/v2\/posts\/82402\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/neclink.com\/index.php\/wp-json\/wp\/v2\/media\/82403"}],"wp:attachment":[{"href":"https:\/\/neclink.com\/index.php\/wp-json\/wp\/v2\/media?parent=82402"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/neclink.com\/index.php\/wp-json\/wp\/v2\/categories?post=82402"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/neclink.com\/index.php\/wp-json\/wp\/v2\/tags?post=82402"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}