目标检测是计算机视觉领域中的重要任务,旨在识别图像或视频中的特定类别物体,并确定它们的位置。与图像分类任务只需判断整个图像属于哪个类别不同,目标检测还需要标记出目标在图像中的边界框。例如在自动驾驶场景中,不仅需要检测道路图像是否包含车辆、人行道和行人,还需要确定它们在图像中的位置。
1 创建目标检测数据集
2 区域提议
区域提议(Region Proposal)是目标检测中的一项重要技术,用于生成可能包含目标物体的候选区域。利用 SelectiveSearch生成区域提议:
3 生成区域提议
4 交并比概念
参考代码:
import selectivesearch from skimage.segmentation import felzenszwalb import cv2 from matplotlib import pyplot as plt import numpy as np import matplotlib.patches as mpatches
img_r = cv2.imread('18.png')img = cv2.cvtColor(img_r, cv2.COLOR_BGR2GRAY)
segments_fz = felzenszwalb(img, scale=200)
plt.figure(figsize=(10,10))plt.subplot(121) plt.imshow(cv2.cvtColor(img_r, cv2.COLOR_BGR2RGB))plt.title('Original Image')plt.subplot(122) plt.imshow(segments_fz)plt.title('Image post \nfelzenszwalb segmentation')plt.show()
def extract_candidates(img): """ 使用选择性搜索(Selective Search)从图像中提取可能包含物体的候选矩形框 参数: img: 彩色图像(BGR格式) 返回: candidates: 列表,每个元素为 [x, y, w, h] 形式的候选框坐标和尺寸 """ img_lbl, regions = selectivesearch.selective_search(img, scale=200, min_size=2000)
img_area = np.prod(img.shape[:2]) candidates = [] for r in regions: if r['rect'] in candidates: continue if r['size'] < (0.05 * img_area): continue if r['size'] > (1 * img_area): continue x, y, w, h = r['rect'] candidates.append([x, y, w, h]) return candidates
img = cv2.imread('18.png')candidates = extract_candidates(img)
fig, ax = plt.subplots(ncols=1, nrows=1, figsize=(6, 6))ax.imshow(cv2.cvtColor(img, cv2.COLOR_BGR2RGB))for x, y, w, h in candidates: rect = mpatches.Rectangle( (x, y), w, h, fill=False, edgecolor='red', linewidth=1 ) ax.add_patch(rect)plt.show()
def get_iou(boxA, boxB, epsilon=1e-5): """ 计算两个矩形框的交并比(Intersection over Union, IoU) 参数: boxA, boxB: 矩形框,格式为 (x1, y1, x2, y2),即左上角(x1,y1)和右下角(x2,y2)的坐标 epsilon: 极小值,防止分母为零 返回: iou: [0,1] 之间的浮点数,越大表示两个框重叠程度越高 """ x1 = max(boxA[0], boxB[0]) y1 = max(boxA[1], boxB[1]) x2 = min(boxA[2], boxB[2]) y2 = min(boxA[3], boxB[3])
width = (x2 - x1) height = (y2 - y1)
if (width 0) or (height 0): return 0.0 area_overlap = width * height
area_a = (boxA[2] - boxA[0]) * (boxA[3] - boxA[1]) area_b = (boxB[2] - boxB[0]) * (boxB[3] - boxB[1])
area_combined = area_a + area_b - area_overlap
iou = area_overlap / (area_combined + epsilon) return iou