社区所有版块导航
Python
python开源   Django   Python   DjangoApp   pycharm  
DATA
docker   Elasticsearch  
aigc
aigc   chatgpt  
WEB开发
linux   MongoDB   Redis   DATABASE   NGINX   其他Web框架   web工具   zookeeper   tornado   NoSql   Bootstrap   js   peewee   Git   bottle   IE   MQ   Jquery  
机器学习
机器学习算法  
Python88.com
反馈   公告   社区推广  
产品
短视频  
印度
印度  
Py学习  »  机器学习算法

论文推荐 | 利用多源数据与机器学习揭示数字隐形群体

北京城市实验室BCL • 昨天 • 17 次点击  
导读

本期为大家推荐的内容为论文《Revealing digitally invisible groups through a machine learning approach using multi-source data》(利用多源数据与机器学习揭示数字隐形群体),发表在 Transactions in Urban Data, Science, and Technology 期刊,欢迎大家学习与交流。


大数据已经成为城市规划决策的重要工具,但数据可得性与代表性在不同空间、时间和社会群体之间并不均衡,可能使部分群体在数据驱动决策中持续“不可见”。研究以发展中国家的土地利用分类为例,将卫星影像、夜间灯光和建筑轮廓等传统地理空间数据,与带地理标记的Twitter帖子、街景影像等数字生成数据结合,利用随机森林开展逐步数据整合,并通过不同数据组合下各类别性能的变化识别代表性缺口。结果显示,非正规住区在Twitter数据中代表不足,难以进入的社区也较少被街景影像覆盖;依赖单一数据源可能强化偏差,而系统评估指导下的互补数据整合能够部分缓解这些缺口。作者进一步建议通过针对性一手调查与参与式制图补充持续存在的数据盲区。


图片

论文相关

图片

题目:Revealing digitally invisible groups through a machine learning approach using multi-source data

利用多源数据与机器学习揭示数字隐形群体

作者:

Wenlan Zhang, Chen Zhong*, Faith Taylor, Yan Liu, Mark Pelling

发表刊物:

Transactions in Urban Data, Science, and Technology

URL:

https://doi.org/10.1177/27541231261457648


摘要ABSTRACT

Big data has emerged as a critical instrument for urban planning and development decision-making. However, the reliability and representativeness of big data constrain its utility. Availability of big data varies significantly across different space, time and socio-demographic groups, particularly in the Global South. This leads to the existence of digitally invisible groups – those who cannot contribute to and benefit from digital data-informed decisions – resulting in the deepening of existing inequalities and further marginalising those already excluded populations. This study presents an example application using land use classification with data from different sources in a developing country context, to explore how certain community groups may be systematically underrepresented or overlooked in specific data and applications. We combine traditional geospatial data (satellite imagery, nighttime light imagery, building footprints) with large-scale, digitally generated data sources (geotagged Twitter posts, street view imagery), and apply a stepwise data integration approach using a random forest classifier. We focus on class-specific changes in performance to infer patterns of uneven data representation. By comparing model outputs across different data combinations, we assess how the inclusion or exclusion of specific datasets influences classification performance. Results indicate that informal settlement areas are underrepresented in geotagged Twitter data, and inaccessible neighbourhoods are poorly captured by street view imagery. Our findings show that reliance on a single data source can reinforce biases, while integrating complementary datasets can partially mitigate these gaps when guided by systematic evaluation. We recommend targeted primary data collection and participatory mapping to address persistent blind spots and improve the inclusiveness of data-informed urban decision-making.

图片

论文展示

图片

更多内容,请点击微信下方菜单即可查询。
请搜索微信号“Beijingcitylab”关注。


图片

Email:BeijingCityLab@gmail.com

Emaillist: BCL@freelist.org

新浪微博:北京城市实验室BCL

微信号:beijingcitylab

网址: http://www.beijingcitylab.org

责任编辑:张业成、黄子沐


Python社区是高质量的Python/Django开发社区
本文地址:http://www.python88.com/topic/200627