MIT與Qatar科學家打造AI系統,辨識假新聞從來源著手
MIT與Qatar科學家打造AI系統,辨識假新聞從來源著手
News from: iThome & MIT Computer Science & Artificial Intelligence Lab
依靠人工查核新聞事實,不僅工作負擔龐大且緩不濟急,麻省理工學院CSAIL及QCRI的科學家合作,基於機器學習開發一套AI系統,能根據來源辨別假新聞,目前已有一定的準確度。
Web site: https://www.csail.mit.edu/news/detecting-fake-news-its-source
麻省理工學院(MIT)電腦科學暨人工智慧實驗室(Computer Science and Artificial Intelligence Lab,CSAIL)與卡塔爾計算研究所(Qatar Computing Research Institute,QCRI)正在打造一基於機器學習AI系統,能夠根據新聞的來源來辨識假新聞,目前在偵測事實可靠性上有65%的準確度,對於政治傾向的判斷則有70%的準確度,將可搭配事實查核網站使用,以減少假新聞流竄的時間。
科學家表示,近來的事實查核世界已產生了危機,諸如Politifact或Snopes等網站都是針對某一主題進行查核,只不過,當他們查明或揭穿事實時,這些假新聞或謠言可能已經繞了地球一圈。
於是CSAIL及QCRI認為最好的作法是不只專注於個別新聞的查證,還必須注重新聞來源,使得他們決定打造一個機器學習系統,可用來快速判斷新聞來源是否可靠或者帶有政治上的偏見。
該研究的主要作者Ramy Baly表示,如果一個網站曾出版過假新聞,那它很有可能會再犯,該系統可自動抓取這些網站的數據,並在第一時間察覺,而且只需要150篇文章來判斷新聞來源的可靠度,之後即可在假新聞被廣泛散布前就逮住它。
這群科學家先取用了來自Media Bias/Fact Check (MBFC)的數據,這是一個集結人類事實查核員的網站,已分析超過2,000個新聞網站的可靠度,從知名的MSNBC到內容農場,並把這些數據匯入機器學習演算法,以建立同樣的分類模式。
該系統還會檢查新聞出處的維基百科頁面來評估新聞來源的可靠程度,例如若維基百科的內容愈豐富,那麼該站就相對可靠,假設出現了「極端」或「陰謀論」等文字,代表它可能有所偏頗。
當把一個新聞出處輸入該系統時,在辨識該出處的新聞真實性(高、中、低)時已有65%的準確度,判斷它是左派、右派或中立立場的準確度則有70%。該研究的共同作者Preslav Nakov表示,目前此一系統還不夠完善,在準確度上仍需加強,最好的方式是與傳統的事實查核機制一起運作。
科學家們也已建立一個內含1,000個新聞來源的開源資料集,並加註了這些新聞來源的真實性及偏見分數,下一步將嘗試把以英文訓練的系統轉為其它語言,也會新增諸如宗教等政治以外的偏見指數。
-------------------------------------------------------------------------

於是CSAIL及QCRI認為最好的作法是不只專注於個別新聞的查證,還必須注重新聞來源,使得他們決定打造一個機器學習系統,可用來快速判斷新聞來源是否可靠或者帶有政治上的偏見。
該研究的主要作者Ramy Baly表示,如果一個網站曾出版過假新聞,那它很有可能會再犯,該系統可自動抓取這些網站的數據,並在第一時間察覺,而且只需要150篇文章來判斷新聞來源的可靠度,之後即可在假新聞被廣泛散布前就逮住它。
這群科學家先取用了來自Media Bias/Fact Check (MBFC)的數據,這是一個集結人類事實查核員的網站,已分析超過2,000個新聞網站的可靠度,從知名的MSNBC到內容農場,並把這些數據匯入機器學習演算法,以建立同樣的分類模式。
該系統還會檢查新聞出處的維基百科頁面來評估新聞來源的可靠程度,例如若維基百科的內容愈豐富,那麼該站就相對可靠,假設出現了「極端」或「陰謀論」等文字,代表它可能有所偏頗。
當把一個新聞出處輸入該系統時,在辨識該出處的新聞真實性(高、中、低)時已有65%的準確度,判斷它是左派、右派或中立立場的準確度則有70%。該研究的共同作者Preslav Nakov表示,目前此一系統還不夠完善,在準確度上仍需加強,最好的方式是與傳統的事實查核機制一起運作。
科學家們也已建立一個內含1,000個新聞來源的開源資料集,並加註了這些新聞來源的真實性及偏見分數,下一步將嘗試把以英文訓練的系統轉為其它語言,也會新增諸如宗教等政治以外的偏見指數。
-------------------------------------------------------------------------

Lately the fact-checking world has been in a bit of a crisis. Sites like Politifact and Snopes have traditionally focused on specific claims, which is admirable but tedious; by the time they’ve gotten through verifying or debunking a fact, there’s a good chance it’s already traveled across the globe and back again.
Social media companies have also had mixed results limiting the spread of propaganda and misinformation. Facebook plans to have 20,000 human moderators by the end of the year, and is putting significant resources into developing its own fake-news-detecting algorithms.
Researchers from MIT’s Computer Science and Artificial Intelligence Lab (CSAIL) and the Qatar Computing Research Institute (QCRI) believe that the best approach is to focus not only on individual claims, but on the news sources themselves. Using this tack, they’ve demonstrated a new system that uses machine learning to determine if a source is accurate or politically biased.
“If a website has published fake news before, there’s a good chance they’ll do it again,” says postdoc Ramy Baly, the lead author on a new paper about the system. “By automatically scraping data about these sites, the hope is that our system can help figure out which ones are likely to do it in the first place.”
Baly says the system needs only about 150 articles to reliably detect if a news source can be trusted — meaning that an approach like theirs could be used to help stamp out new fake-news outlets before the stories spread too widely.
The system is a collaboration between computer scientists at MIT CSAIL and QCRI, which is part of the Hamad Bin Khalifa University in Qatar. Researchers first took data from Media Bias/Fact Check (MBFC), a website with human fact-checkers who analyze the accuracy and biases of more than 2,000 news sites; from MSNBC and Fox News; and from low-traffic content farms.
They then fed those data to a machine learning algorithm, and programmed it to classify news sites the same way as MBFC. When given a new news outlet, the system was then 65 percent accurate at detecting whether it has a high, low or medium level of factuality, and roughly 70 percent accurate at detecting if it is left-leaning, right-leaning, or moderate.
The team determined that the most reliable ways to detect both fake news and biased reporting were to look at the common linguistic features across the source’s stories, including sentiment, complexity, and structure.
For example, fake-news outlets were found to be more likely to use language that is hyperbolic, subjective, and emotional. In terms of bias, left-leaning outlets were more likely to have language that related to concepts of harm/care and fairness/reciprocity, compared to other qualities such as loyalty, authority, and sanctity. (These qualities represent a popular theory — that there are five major moral foundations — in social psychology.)
Co-author Preslav Nakov, a senior scientist at QCRI, says that the system also found correlations with an outlet’s Wikipedia page, which it assessed for general — longer is more credible — as well as target words such as “extreme” or “conspiracy theory.” It even found correlations with the text structure of a source’s URLs: Those that had lots of special characters and complicated subdirectories, for example, were associated with less reliable sources.
“Since it is much easier to obtain ground truth on sources [than on articles], this method is able to provide direct and accurate predictions regarding the type of content distributed by these sources,” says Sibel Adali, a professor of computer science at Rensselaer Polytechnic Institute who was not involved in the project.
Nakov is quick to caution that the system is still a work in progress, and that, even with improvements in accuracy, it would work best in conjunction with traditional fact-checkers.
Nakov is quick to caution that the system is still a work in progress, and that, even with improvements in accuracy, it would work best in conjunction with traditional fact-checkers.
“If outlets report differently on a particular topic, a site like Politifact could instantly look at our fake news scores for those outlets to determine how much validity to give to different perspectives,” says Nakov.
Baly and Nakov co-wrote the new paper with MIT Senior Research Scientist James Glass alongside graduate students Dimitar Alexandrov and Georgi Karadzhov of Sofia University. The team will present the work later this month at the 2018 Empirical Methods in Natural Language Processing (EMNLP) conference in Brussels, Belgium.
The researchers also created a new open-source dataset of more than 1,000 news sources, annotated with factuality and bias scores, that is the world’s largest database of its kind. As next steps, the team will be exploring whether the English-trained system can be adapted to other languages, as well as to go beyond the traditional left/right bias to explore region-specific biases (like the Muslim world’s division between religious and secular).
“This direction of research can shed light on what untrustworthy websites look like and the kind of content they tend to share, which would be very useful for both web designers and the wider public,” says Andreas Vlachos, a senior lecturer at the University of Cambridge who was not involved in the project.
Nakov says that QCRI also has plans to roll out an app that helps users step out of their political bubbles, responding to specific news items by offering users a collection of articles that span the political spectrum.
“It’s interesting to think about new ways to present the news to people,” says Nakov. “Tools like this could help people give a bit more thought to issues and explore other perspectives that they might not have otherwise considered."



留言
張貼留言