大家好,又见面了,我是全栈君,今天给大家准备了Idea注册码。
在pig中, dump和store会分别完毕两个MR, 不会一起进行
1:载入名用正則表達式:
LOAD ‘/user/wizad/data/wizad/raw/2014-0{6,7-0,7-1,7-2,7-3,8}*/3_1/adwords*’
2:filter的几种简单使用方法:
按值过滤
FILTER clickDate_all BY log_type==’2′;
FILTER mapping_table BY mapping_ad_network_id==’3′ AND mapping_type==’5′;
test =FILTER allRow BY (ad_id==’14997′ OR ad_id==’14998′ OR ad_id==’14999′) AND log_type==2;
test=FILTER allRow BY (INDEXOF(ad_id,’14997′)==0 OR INDEXOF(ad_id,’14998′)==0 OR INDEXOF(ad_id,’14999′)==0) AND log_type==2;
配合size函数
FILTER count_imei BY (SIZE(cimei)>14 AND SIZE(cimei)<17);
正則表達式
FILTER cimei2 BY NOT cimei MATCHES ‘^[0-9]*$’;
FILTER cmac2 BY cmac MATCHES ‘/[A-F\d]{2}:[A-F\d]{2}:[A-F\d]{2}:[A-F\d]{2}:[A-F\d]{2}:[A-F\d]{2}/’;
3:排序
ORDER province_count BY $2 DESC;
4:CONCAT函数的使用。可用于生成独立的一列,如count了的一个数,前面加一列名称
FOREACH origin_cleaned_data GENERATE CONCAT(‘<-_’,’->’) AS cou,guid,log_type;
read_social_14 =FOREACH metadata_social_14 GENERATE CONCAT(’14’,’==’),guid_social;
all_id =FOREACH allRow GENERATE id,CONCAT(‘_’,’-‘) as cc;
5:过滤空值,将空值改成取值unknown。
条件表达式“(推断式)?a:b”的应用:直接对列操作
origin_historical = FOREACH origin_cleaned_data GENERATE wizad_ad_id,guid,log_type,
((province_region_id == ”) ? ‘unknown’ : province_region_id)
6:切分成不同子集,按值:
SPLIT geelyTuiGuang INTO android IF os_id==1,ios IF os_id==2;
SPLIT ios INTO ios6 IF (INDEXOF(os_version,’7′)!=0),ios7 IF INDEXOF(os_version,’7′)==0;
SPLIT allCleaned INTO log_42 IF (
((chararray)$34==’1′ OR (chararray)$34==’2′ OR (chararray)$34==’3′ OR (chararray)$34==’1′ OR (chararray)$34==’4′)
AND
(INDEXOF((chararray)$35,’.’)>0)
AND
((chararray)$36==’1′ OR (chararray)$36==”)
),
log_43 IF (
((chararray)$34==’1′ OR (chararray)$34==’2′)
AND
((chararray)$35==’1′ OR (chararray)$35==’2′ OR (chararray)$35==’3′ OR (chararray)$35==’1′ OR (chararray)$35==’4′)
AND
(INDEXOF((chararray)$36,’.’)>0)
);
7:replace函数替换值
FOREACH ios6 GENERATE imei,mac_address as cmac,REPLACE(idfa,’null’,”);
8:数据流过滤
en_guid =STREAM duimei THROUGH `awk -F”,” ‘{if($3 == “null”) print $1″,”$2″,”; else print $0}’`;
9:强制转换:
cleaned_data_42 =FOREACH log_42 GENERATE
(chararray)$1 AS wizad_ad_id:chararray,
(chararray)$2 AS guid:chararray,
(chararray)$6 AS log_type:chararray,
(chararray)$18 AS imei:chararray,
(chararray)$22 AS idfa:chararray,
(chararray)$23 AS mac_address:chararray
10内置函数REGEX_EXTRACT,使用正則表達式:
allAdId =FOREACH allRow GENERATE REGEX_EXTRACT((chararray)$3,'(.*) (.*)’,1) AS time,REGEX_EXTRACT((chararray)$0,'(.*)_(.*)’,1) AS adn,$6 AS ad_id;
allAdId =FOREACH allRow GENERATE REGEX_EXTRACT(create_time,'(.*) (.*)’,1) AS time,ad_id;
发布者:全栈程序员-用户IM,转载请注明出处:https://javaforall.cn/117975.html原文链接:https://javaforall.cn
【正版授权,激活自己账号】: Jetbrains全家桶Ide使用,1年售后保障,每天仅需1毛
【官方授权 正版激活】: 官方授权 正版激活 支持Jetbrains家族下所有IDE 使用个人JB账号...