Affichage des articles dont le libellé est Analyse des données. Afficher tous les articles
Affichage des articles dont le libellé est Analyse des données. Afficher tous les articles

vendredi 28 octobre 2011

Jubatus : analyse en temps réel de gros volumes de données

De nombreux systèmes d'analyse de données sont basés sur un traitement séquentiel (batch). Cependant, ce type de traitement n'est pas assez efficace dans le cas d'applications nécessitant une analyse temps-réel sur un grand nombre de données. Le traitement séquentiel impose à un serveur d'attendre que toutes les données précédemment reçues soient traitées, ce qui augmente peu à peu le temps de réaction et ne permet donc pas de répondre à des situations où la spontanéité est un facteur critique.

L'un des leaders japonais des technologies de l'information, NTT, a développé conjointement avec l'entreprise japonaise Preferred Infrastructure une technologie appelée Jubatus permettant une analyse en temps-réel d'un gros volume de données.

Le nom Jubatus provient du nom scientifique du guépard....

Lire la suite >>

http://www.bulletins-electroniques.com/actualites/68051.htm

dimanche 7 août 2011

Italian academia is a family business, statistical analysis reveals


Frequency of last names in disciplines, institutions suggests rampant nepotism


Unusually high clustering of last names within Italian academic institutions and disciplines indicates widespread nepotism in the country's schools, according to a new computational analysis.

By comparing the frequency of last names among more than 61,000 professors in medicine, engineering, law, and other fields, University of Chicago researcher Stefano Allesina found the pattern to be incompatible with unbiased, equal opportunity hiring. The analysis, published online in the journal PLoS ONE, refutes the notion that recently publicized cases of academic nepotism in Italy were isolated incidents.

"It's not a few bad apples, it's really bad," said Allesina, PhD, assistant professor of ecology & evolution. "I found that in many disciplines there are much fewer names than you would expect to find at random, indicating a very, very high probability of nepotistic hires."

In recent years, several scandals have hit Italian academia over the hiring of close family members to prominent faculty positions at public universities. At the University of Bari, nine relatives from three generations of a single family are on the economics faculty, several newspapers reported last year. The chancellor of Sapienza University in Rome was recently investigated by an Italian news program after the hiring of his wife, son, and daughter to medical faculty positions.

To measure the full magnitude of nepotism in Italian academia, Allesina turned to a public database created by the Italian Ministry of Education. Included was first and last name information for over 61,000 tenured professors from 94 institutions, along with their department and sub-discipline.

Allesina used the pool of last names to run a simple analysis of name frequency. More than 27,000 different last names appeared at least once in the dataset, and Allesina sought to test whether certain names appeared more often than expected in a given field. So he programmed the computer to conduct one million random drawings from the pool of names to see how probable it was to obtain the number of last names that exist in the real-life data.

For example, of the 10,783 faculty members working in medicine, 7,471 distinct last names were found. But in one million random drawings from the full pool of names, Allesina's program never came up with fewer than 7,471 unique names, indicating an improbable frequency of last names indicative of nepotistic hiring.

"It's very basic, anybody with a laptop can do this analysis," Allesina said. "I wanted to keep it as coarse-grained and simple as possible. Because then it's more powerful – if this works, anything else will work. Even this very simplistic analysis can find that some disciplines are above and beyond what one could expect."

Allesina repeated the computation for 28 academic fields, finding the highest likelihood of nepotism in industrial engineering, law, medicine, geography, and pedagogy. Fields with the distribution of names closest to random – and thus with the lowest likelihood of nepotism – were linguistics, demography, and psychology.

In another analysis, Allesina looked at the geographic distribution of nepotism across Italy. In this model, he tested the probability of sharing a last name with another faculty member in the same field of study, and mapped those probabilities from the north to south of the country. The model discovered a stark north-to-south gradient, with the probability of nepotism increasing as one looked south, peaking on the island of Sicily.

The distribution mirrors similar negative statistics such as infant mortality, organized crime, and suicide rate that are higher in southern Italy compared to northern regions.

"For an Italian, this is not that surprising," Allesina said. "It is a narrative of two separate countries, where in the public sector we have more problems in the south."

The research suggests that nepotism is a pervasive problem in Italian academia, a blemish that undercuts the quality of advanced education in the country and drives professionals abroad when they fail to find job opportunities at home.

"In Italy, there is an enormous brain drain," Allesina said. "I think these kind of hiring practices contribute a lot to this brain drain and the fact that Italian universities are not ranked very high internationally."

The Italian government passed a university reform law in late 2010 to establish new rules for academic hiring and the distribution of grant money despite opposition from professors and students. Allesina said that his analysis could easily be repeated in the future to test whether the reforms truly reduced nepotism in the system, and hopes it will also be applied to other fields and countries where unfair hiring practices are suspected.

"I think this problem with the university is really the tip of the iceberg, a place where it's really apparent what the problems of Italy are," Allesina said. "But I wanted to keep the methods easy enough so that a researcher can do it not just for universities, but for other places in the public sector. The government has all these numbers and names, if they wanted to do it they could."


Public release date: 3-Aug-2011

Contact: Robert Mitchum
robert.mitchum@uchospitals.edu
University of Chicago Medical Center

http://www.eurekalert.org/pub_releases/2011-08/uocm-iai080111.php

dimanche 31 juillet 2011

La cohésion, une mesure fiable pour identifier les communautés


Des chercheurs de l’Inria ont validé par l’expérience leur approche qui vise à mieux identifier les communautés dites recouvrantes. Résultat : leur métrique, la cohésion sociale, tend à être une mesure de qualité.

En février dernier, l’équipe de chercheurs DNET* à l’ENS de Lyon, lançait une expérience, Fellows, pour mettre en application un algorithme permettant de mieux identifier les communautés recouvrantes. Un article vient de paraître pour annoncer les résultats après six mois d’expérimentation. Selon les chercheurs, la pratique a bien validé la théorie. "Nous avons confirmé que notre métrique, la cohésion est une mesure de qualité", explique le professeur Eric Fleury, directeur de l’équipe DNET. Leur algorithme de génération de groupe, selon le chercheur, permet de dépasser les modèles actuels d’identification de communautés, et d’être plus réaliste. Dans leur idée, il pourrait par exemple servir à mieux organiser ses groupes dans Facebook, ou ses cercles dans Google +.

Des triangles pour mieux capturer la cohésion

Sur Facebook, par exemple, tous les liens entre individus sont similaires : tout le monde est "ami". L’approche manque de granularité. "Dans les graphes plus standards, un noeud (un ami) appartient en général à une seule communauté. Ce n’est pas conforme à la réalité : on sait très bien que dans la vraie vie, des noeuds peuvent se recouvrir. Ces derniers sont d’ailleurs importants, puisque ce sont ceux qui font passer une information d’une communauté à l’autre", explique Eric Fleury. L’équipe a eu l’idée de s’inspirer de la sociologie et de la notion fondamentale de triangle, et de l’appliquer à l’analyse des réseaux, pour arriver à une identification de communautés plus juste. "Le principe n’a rien de révolutionnaire, et est très simple : si deux personnes sont amies avec une troisième, elles sont amies entre elles", reprend Eric. Les chercheurs exploitent les informations "qui est ami avec qui" sous la forme d’une mesure. C’est la cohésion, qui est la pierre angulaire de leur algorithme, et qui quantifie à quel point un ensemble de personnes forment un groupe social.

L’expérience a validé le modèle

Fellows a permis de confronter la théorie à la réalité et prouver que la cohésion était une approche valable. L’expérience a été menée à large échelle sur Facebook. L’utilisateur s’y connecte et valide l’accès à son compte Facebook (via une API). Le programme présente ensuite à l’utilisateur la liste des groupes proposés automatiquement et demande de les noter, de une à quatre étoiles selon la pertinence. Par ailleurs, Fellows propose — si l’utilisateur le souhaite — de créer automatiquement des listes d’amis dans Facebook, réduisant ainsi le temps qu’il lui est nécessaire pour le faire lui-même. Selon l’article, en six mois, les 2157 participants ont amené à la détection de 67750 groupes, et 43589 ont reçu une note. 25,1% ont eu une étoile, 21,8 % deux étoiles, 22,5 % ont reçu trois étoiles, et 30,7 % ont reçu quatre étoiles. Le but n’était pas d’obtenir la plus haute proportion de 4 étoiles, mais bien de vérifier la pertinence de leur modèle. "Nous avons pu prouver qu’il existe une corrélation entre la cohésion, et les notes données à chaque groupe. Ceci prouve que notre algorithme fait sens", conclut Eric Fleury. Les prochains travaux de l’équipe porteront sur l’adaptation de ce modèle aux réseaux complexes.

* L'equipe DNET est commune à l'institut et à l'École Normale Supérieure de Lyon, au sein du laboratoire LIP (Laboratoire de l'Informatique du Parallélisme). Elle conduit des recherches théoriques et expérimentales sur les réseaux sociaux afin de mieux appréhender leur structuration et la dynamique des processus de diffusion d'information en leur sein. Elle est dirigée par Eric Fleury, professeur ENS Lyon. Les travaux de recherche liés à fellows et l'expérimentation associée sont menés par Adrien Friggeri, doctorant dans l'équipe DNET.


Publié le 29 juillet 2011

http://www.atelier.net/articles/cohesion-une-mesure-fiable-identifier-communautes

vendredi 29 juillet 2011

New Geographic Data Analysis Gives Historians a Futuristic Window Into the Past

"Spatial humanities," the future of history   

 
Gettysburg Battlefield Visitors to the Gettysburg battlefield will see some undeveloped vistas like this one, but they won’t see what Robert E. Lee saw, because time has altered the landscape. New analysis with geographic information software has allowed historians to virtually reconstruct what the scene looked like in July 1863.Wikimedia Commons
 
Even using the most detailed sources, studying history often requires a great imagination, so historians can visualize what the past looked and felt like. Now, new computer-assisted data analysis can help them really see it.

Geographic Information Systems, which can analyze information related to a physical location, are helping historians and geographers study past landscapes like Gettysburg, reconstructing what Robert E. Lee would have seen from Seminary Ridge. Researchers are studying the parched farmlands of the 1930s Dust Bowl, and even reconstructing scenes from Shakespeare’s 17th-century London.

But far from simply adding layers of complexity to historical study, GIS-enhanced landscape analysis is leading to new findings, the New York Times reports. Historians studying the Battle of Gettysburg have shed light on the tactical decisions that led to the turning point in the Civil War. And others examining records from the Dust Bowl era have found that extensive and irresponsible land use was not necessarily to blame for the disaster.

GIS has long been used by city planners who want to record changes to the landscape over time. And interactive map technology like Google Maps has led to several new discoveries. But by analyzing data that describes the physical attributes of a place, historians are finding answers to new questions.
Anne Kelly Knowles and colleagues at Middlebury College in Vermont culled information from historical maps, military documents explaining troop positions, and even paintings to reconstruct the Gettysburg battlefield. The researchers were able to explain what Robert E. Lee could and could not see from his vantage points at the Lutheran seminary and on Seminary Hill. He probably could not see the Union forces amassing on the eastern side of the battlefield, which helps explain some of his tactical decisions, Knowles said.

Geoff Cunfer at the University of Saskatchewan studied a trove of data from all 208 affected counties in Texas, New Mexico, Colorado, Oklahoma and Kansas — annual precipitation reports, wind direction, agricultural censuses and other data that would have been impossible to sift through without the help of a computer. He learned dust storms were common throughout the 19th century, and that areas that saw nary a tiller blade suffered just as much.

The new data-mapping phenomenon is known as spatial humanities, the Times reports. Check out their story to find out how advanced technology is the future of history.

[New York Times]

http://www.popsci.com/technology/article/2011-07/high-tech-geographic-data-analysis-helps-historians-see-past